Compare commits

13 Commits

Author SHA1 Message Date
MechaCat02
053aa81953 Cutscene voice: resolve movies via continuous-stream cue regions
Cracked how the game maps a movie to its voice track and reimplemented it,
fixing wrong/short voices for RT and hokyu cutscenes.

## Mechanism (RE'd from a Canary file-I/O trace, user-confirmed by ear)

The movie voices are ONE continuous XMA stream, physically chunked into
`<lang>\Movie\VOICE_*.slb` (and `\etc\VOICE_D_*.slb`) sound.pak TOC entries
whose boundaries do NOT align with the cutscene cues. A single cue routinely
spans two `.slb` chunks, so a `.slb` does not necessarily hold the track its
name claims (they are off-by-one vs the cues). Each cue ends at an inline
trailer descriptor `(sound_id:u32be, 0x11, …)` whose id repeats at +0x800.

Resolution chain (all derivable on-disc, no guest-code RE):
  1. movie -> cue token         (ADVERTISE_MOVIE manifest VOICETRACK)
  2. cue token -> sound-id       (master sound registry, tables.pak
                                  a2c8c185=eng / b04238e0=jpn, schema 0x13cb84ba)
  3. sound-id -> [start,end)      scan the stream for trailer(id) and its
                                  predecessor; cue audio = the bytes between,
                                  read continuously across chunk boundaries.

Validated: decoded voice length matches each movie within ~0.4s, including
cross-boundary cues (ADV/RT01A/RT01B/RT01C/S00A/S01A + the 5 bound hokyu).

## Changes

- sylpheed-formats/src/movie_voice.rs (new): registry_voice_ids(),
  find_descriptor(), find_descriptor_before() (gap-tolerant predecessor).
  Unit-tested.
- viewer iso_loader.rs: resolve_movie_voice_region() + decode_voice_region()
  (continuous read via existing read_segment_range + to_xma_riffs). Voice/
  audio requests now region-first; a `\Movie\` token that fails region stays
  SILENT (never plays the wrong off-by-one chunk). Removed the old
  movie_is_demo_based suppress hack.
- Anchor tries subdirs Movie/etc/Voice so hokyu `\etc\VOICE_D_*` resolve too.
- Examples resolve_all/validate_cues/hokyu_final document + validate the pipeline.

## Status

- RT / S / ADV cutscene voices: CORRECT (user-confirmed).
- 5 manifest-bound hokyu (VOICE_D_450-454): CORRECT (user-confirmed).

## OPEN (handoff)

13 of 18 hokyu movies have NO manifest VOICETRACK; the game reuses the 5
recordings across stages via a code rule. `hokyu_fallback_token()` currently
guesses by category (ship LS/DS x source carrier-A/tanker-H: LS/A->450,
LS/H->453, DS/A->452, DS/H->454). USER REPORTS THIS IS SOMETIMES WRONG.
- LS/carrier has two takes (450 from s02A, 451 from s09A); the early-vs-late
  split is unknown — unbound LS/A currently all default to 450.
- Definitive fix = Canary `--phase_a_fileio_only` file.read trace of an unbound
  hokyu (e.g. hokyu_LS_s03A) to see which VOICE_D the game actually streams,
  then encode the real rule. Same method that cracked the RT mapping.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-23 06:53:55 +02:00
MechaCat02
ed5c2f24c3 feat(formats): decode the movie manifest's full mission/phase structure
The movie manifest (tables.pak, schema 0x067025b9) is the game's authoritative mission -> phase -> cutscene table. Previously we only scraped VOICE_/SUBTITLE_/.wmv strings by prefix; now we decode the real two-array (slot-keys, then value-records) structure.

MovieEntry gains slot, kind (System/Intro/Phase/PhaseEnd/Supply), mission, phase, subtitle, and telop (the previously-ignored .prt on-screen text overlay). Valueless slots (stages with no resupply video) are recovered without decoding the binary node region, via the fully-derivable MS<NN><X> -> S<NN><X> anchors.

Backward compatible: movie/voice_token and resolve_voice_entry/voice_token/is_manifest are unchanged. Validated end-to-end on the real disc: 101 entries (104 slots - 3 gaps), typed System=6/Intro=27/Phase=32/PhaseEnd=18/Supply=18. examples/manifest_map.rs dumps the full map.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 22:49:19 +02:00
MechaCat02
eee1607e1b feat(formats,viewer): standalone voice player covers all line categories
list_voice_clips now enumerates every spoken-line category so the standalone audio player covers them all — bound movie voices (\Movie\), in-mission radio (\Voice\, \etc\), and mission briefings (\Briefing\, whose BR<NN>_<MM> names lack "VOICE"); music (BGM_*) stays excluded.

Viewer voice-library fixes: invalidate the library when a new game source loads, and don't latch an empty result as "loaded" (sounds.tbl always has thousands of clips, so empty = a failed read) so re-opening the menu retries.

Voice decode reverted to the first sub-wave (base behaviour): a naive multi-sub-wave concat produced the wrong track for RT movies. The per-sub-wave splitter (to_xma_riffs) is kept + documented for the eventual real fix, which needs ground-truth playback instrumentation to resolve segments-vs-takes-vs-timecode placement.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-21 22:49:09 +02:00
MechaCat02
95de29f290 feat(formats,viewer): movie subtitles, voice decode, and the movie manifest
The movie cutscene subtitle + voice pipeline, driven by the ADVERTISE_MOVIE
manifest (the authoritative movie -> subtitle -> voice index). Also flushes
several sessions of local WIP (async viewer loading, grouped-pool XBG7/hero-ship
decode, drawlog tooling). See docs/HANDOFF-movie-voice-subtitles-2026-07-19.md.

Subtitles (movie_subtitle.rs):
- Full movie->track->text chain; join multi-line captions sharing one timing
  (fixes S13A dropped "Look at it father" line); overlap-safe active_cues();
  Latin-1 accents preserved.

Voice (slb.rs): XACT .slb -> XMA1 RIFF; take the FIRST sub-wave bounded by its
declared data size (fixes S10-S16 alternate-take garble); list_voice_clips.

Manifest (movie_manifest.rs): parse ADVERTISE_MOVIE (0x5B983A08) for the real
movie->voice binding (not always VOICE_<movie>; e.g. hokyu -> VOICE_D_* in etc\).
Resolve the token's sound.pak path via sounds.tbl. DIRECT bindings only — the
demo-id shared-clip fallback for unbound hokyu movies was verified WRONG in-game
and reverted (unbound hokyu stay unvoiced; correct join key is an OPEN problem).

Viewer: manifest-driven voice (movie player toggle + solo button), standalone
"Voice Lines" browser, stacked caption overlay.

Tests: 46 formats-lib + 11 viewer-lib + movie_manifest/movie_subtitle/slb disc
tests; full workspace green.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 20:15:41 +02:00
MechaCat02
578c71a1b9 [formats,viewer] Decode grouped-pool XBG7 models; fix mesh routing + albedo
Revert the mis-guided "triangle STRIP" reading (a937779): XBG7 index
buffers are triangle LISTS (prim=4 in the GPU capture; stored-normal
agreement 1.000 as a list vs ~0.49 as a strip). Reinstate the
winding-consistency gate.

Crack the grouped-pool layout used by the hero ship and ~150 detailed
models. A resource's sub-meshes share one index pool (buffers 4-byte
aligned, in descriptor-marker order) followed by one vertex pool (each
vtx_count*stride, same order); the vertex pool is 4-byte aligned after
the index pool. The whole resource is derived from one anchored pivot:
  ib0 = vb0 - span - pad   (pad in 0..=3, alignment)
  ib[i] = align4(ib[i-1] + idx_count[i-1]*2)
  vb[i] = vb[i-1] + vtx_count[i-1]*stride
vb0 is found via the unit-normal vertex-run scan; the alignment is
confirmed by validating the LARGEST sub-mesh (most reliable), after
which the rest are read/validated. A single marker reduces this to the
existing adjacency anchor, so grouped generalises it.

Results (cross-checked against a Canary GPU draw-log capture of
DeltaSaber_T.xpr):
- DeltaSaber_T f001 = body + 7 detail parts = 8650 tris, every sub-mesh
  0-degenerate / full-coverage / winding-agreement 1.000.
- All 19 previously-declined weapon models now decode (they hit pad 2
  and/or lead with a tiny bracket that broke a markers[0] pivot). Corpus
  audit: 0/146 geometry files fail (was 19 -> flat-texture in the viewer).
- Stages unchanged (single-marker path is byte-identical; grouped falls
  back to the old anchor on failure). 7/7 disc tests green incl. the
  strict stage quality audit; new hero_ship_grouped_pool_decodes test.

Viewer: route by count_xbg7; --only matches an exact model name (so the
neutral pose renders without the mnv*/turn180 animation poses). Fix the
albedo matcher: match a sub-model to its _col map by entity stem
(e007_bdy_01 -> e007_col) instead of a full-name prefix, lifting stage
sub-model texturing from ~25% to ~95% (the rest were flat grey). Colour
correctness (channel order/sRGB) remains a separate dynamic-RE item.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-18 19:54:39 +02:00
MechaCat02
a937779c77 [formats] XBG7 meshes: index buffers are triangle STRIPS, not lists
The weapon/prop models (rou_fxxx_wep_nn.xpr) rendered with triangular holes in
the hull and misplaced/extra spikes. Root cause: XBG7 sub-mesh index buffers are
triangle STRIPS, but the decoder read them as triangle LISTS — a strip of N
indices is N-2 triangles, a list is only N/3, so ~2/3 of every hull was missing
(the holes) and each list-triple of strip data spanned unrelated vertices (the
spikes). The tell: triangles-per-vertex was 0.5-1.0 with 0 unreferenced vertices
(a closed surface needs ~2.0).

- mesh.rs: expand_triangle_strip() converts strip -> list with alternating
  winding, skipping degenerate triangles (repeated index = strip restart).
  Applied in both from_xpr2 (weapons) and read_pool_mesh (stage sub-models).
  All 71 weapon submeshes now have healthy ratios; hulls fill in (wep_03 cannon,
  wep_04 pod, wep_11, wep_37 verified coherent).
- Module doc updated (strip, not list).
- sylpheed-cli: `mesh info` reports per-submesh degenerate/unref/spanning-triangle
  counts; `mesh render` gains --dist (camera zoom). XMESHDBG env dumps the
  descriptor index markers.

Known residual: an index buffer may concatenate several strips with no degenerate
bridge, leaving ~2 spanning "spike" triangles at each restart (index jump to a new
vertex region) — <1% of tris on most models. Decoding the restart mechanism is a
follow-up (a blanket spanning-filter is unsafe: clean models have legit elongated
tip triangles).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-18 13:04:12 +02:00
MechaCat02
7143ea18fe [formats] Fully crack T8aD: decode the header's per-tile offset table
Follow-up to the 256×256-tiling reversal — the residual boundary seams on
≥2×2-tile textures are gone. Decoding the file header revealed the exact
layout (no oracle capture needed):

  0x00  44        base header
  0x1c  4         tile count = ceil(w/256) * ceil(h/256)
  0x2c  tiles*4   BE-u32 absolute offset of each row-major 256×256 tile
  <off> 16        per-tile header, then tile_w*tile_h*4 A8R8G8B8 pixels

The old seams came from ignoring the offset table and the 16-byte per-tile
header (contiguous-packing drifted 16 bytes per tile). Small textures decoded
before only by luck: one tile puts pixels at 44+4+16 = 64, the old type-1
"header size". Now the 8AX title background, the prselect_win1 window frame,
and preff04 all decode pixel-perfect (verified against the title screen).

- t8ad.rs: parse() walks the offset table; dropped detile_256 + the
  header-size-by-type table. Non-tilecount entries (field != ceil*ceil) return
  None (likely DXT/other, deferred). New multi-tile round-trip test.
- sylpheed-cli: pak textures decodes via parse(); XDUMPHDR=1 dumps the base
  header + offset table for RE. Removed the now-obsolete XTILE/XSCORE/XDESTRIP
  experiment knobs and the reconstruct_tiles/TV-sweep helpers.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-18 11:55:25 +02:00
MechaCat02
4ba723b9a5 [formats] Crack large-T8aD tiling: 256×256 row-major raster tiles
Verified against the running game (freeze-fixed Canary title screen) that
T8aD colours are correct (no R↔B swap) and that wide textures were garbled
by an unreversed tile layout, not a colour bug.

Reversed the layout: large T8aD are stored as 256×256 raster tiles in
row-major order, each tile raster internally, edge tiles clipped to the
image (payload is exactly w*h*4, no padding). Surfaces ≤256px wide are a
single tile column, identical to plain linear — which is why small UI
textures always decoded correctly.

- t8ad.rs: add detile_256() + apply in parse(). Single tile-row / single
  tile-column textures now decode exactly (ptcopyright, ptbtn, ptlogo1 =
  clean "PROJECT" logo, previously pure noise). Known residual: ≥2×2-tile
  textures (>256 in both dims, e.g. the 8AX background) come out coherent
  but with boundary seams — the multi-tile order is a subtle swizzle, TBD.
- sylpheed-cli: new `pak textures <pak> <out>` command — decodes every
  T8aD (direct, RATC-nested, LSTA frames) to PNG for A/B against the game.
  Env debug knobs for layout RE: XTILE/XDESTRIP/XDETILE_W/XREINTERPRET_PITCH
  /XSCORE (TV-ranked tile sweep), plus --verbose size classification.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 23:37:35 +02:00
MechaCat02
d23339a3aa feat(viewer): stage thumbnail grid, per-model albedo, camera zoom fix
Route stage containers to stage_models (decode both single + stage paths, keep
whichever yields more geometry) — fixes stages showing the dummy-root "black
cube". Present a stage's sub-models as a thumbnail grid: each recentred and
uniformly scaled to a fixed cell so a 1-unit prop and a 3000-unit structure are
equally visible; skybox-plane / stray-anchor models (>20000u) are culled. Each
sub-model is textured by matching its name to the container's `_col` albedo.

Fix the orbit camera: zoom radius was hard-capped at 50u with fixed clip planes,
trapping the camera inside multi-thousand-unit stages. Zoom limits and near/far
planes now scale to the framed model.

Adds docs/HANDOFF-stage-mesh-2026-07-12.md.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-12 20:56:45 +02:00
MechaCat02
2053f31d17 feat(formats,cli): decode XBG7 stage containers + fix weapon layout
Stage containers (hidden/resource3d/Stage_S*.xpr) are collections of up to
~450 enemy/prop sub-models, not single meshes. Xbg7Model::stage_models decodes
them: each resource is a [12-byte header][index buffer][vertex buffer] block
(index count = descriptor marker, vertex count = u32 32 bytes before it) whose
on-disc offset is NOT stored, so it is located by content — one O(file) pass per
stride finds vertex-buffer starts (unit NORMAL at +12 whose previous slot isn't)
and each resource is pinned to the candidate whose indices validate and produce
non-degenerate, well-connected triangles. A connectivity gate (mean triangle
edge <= 0.28x the bbox diagonal) rejects spiky mis-anchors. ~4993 sub-models
decode across the 22 stages.

The same insight fixes the weapon single-model layout: it is [12-byte header]
[index][vertex] too, not [index][12-byte gap][vertex]. Reading indices from the
block start turned the 12 header bytes into 6 junk indices (2 leading degenerate
triangles — the recurring stray-triangle artifact) and dropped the last 6 real
indices. Skipping the header leaves vertex offsets identical (coverage unchanged
at 36/166) and corrects the triangle list. Verified on wep_00/03/04.

Adds `sylpheed-cli mesh {info,render}` — a headless software rasterizer that
writes a shaded PNG (orthographic, z-buffered, two-sided), so recovered geometry
can be verified without the GUI. Stages render as a normalised thumbnail grid;
--only filters sub-models.

Tests: stage_models_{decode,sweep,quality_audit}; docs/re updated (xbg7-mesh.md
evidence log + INDEX.md).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-12 20:56:35 +02:00
MechaCat02
4096b2d2a5 feat(formats): declaration-driven XBG7 decode (variable stride)
Dynamic-RE follow-up: capture Canary's GPU vertex-fetch + draw calls and
feed the ground truth back into the static decoder.

The GPU capture confirmed the reverse-engineered layout exactly — meshes
draw as triangle LISTs (prim=4) with pos f32x3 @0, normal f16x4 @0x0C,
uv f16x2 @0x14 — and revealed the XBG7 vertex format is NOT fixed-stride:
models omit elements (stride 20 = pos+normal, no UV; 24 = pos+normal+uv).

- mesh.rs: parse the descriptor's vertex declaration ({offset, format-code,
  usage} triples; 0x2A23B9=pos f32x3, 0x1A2360=normal f16x4, 0x2C235F=uv
  f16x2) to drive per-model stride + element offsets, instead of assuming
  stride 24. Coverage 25 -> 36 fully-validated models (e.g. the Stage_S*
  props, which are pos+normal only). Same index-range + unit-normal safety
  gates; complex/mismatched layouts still declined.
- Endianness note (documented): the capture's fetch endian=k8in32 describes
  the GPU's guest-memory copy, NOT the .xpr file bytes — reading the file
  with k8in32 breaks the normals (|n|->1.33); naive big-endian per element
  is correct (the game rearranges vertex data on load).
- tests: a coverage-regression test (>=35 models) + the existing weapon /
  body-declined disc tests still pass.

Confirmed against a Canary draw-log capture; see docs/re/structures/
xbg7-mesh.md (evidence log + declaration table).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-12 18:42:21 +02:00
MechaCat02
ef6e448268 feat(formats,viewer): decode XBG7 meshes + 3D model preview
Reverse-engineer the single-stream XBG7 geometry layout (clean-room: hex
inspection + geometric validation of the retail disc's
hidden/resource3d/*.xpr, no game code copied) and present models as
textured 3D meshes in the explorer.

Format (docs/re/structures/xbg7-mesh.md): XBG7 geometry resources sit
inside XPR2 containers alongside TX2D textures. For ~25 single-stream
models (weapons, simple props) the data section is a sequence of
sub-meshes, each an u16-BE triangle-list index buffer followed (after a
fixed 12-byte header) by a stride-24 vertex buffer whose declaration is
in the descriptor: POSITION f32x3 @0x00, NORMAL f16x4 @0x0C, TEXCOORD
f16x2 @0x14. Sub-mesh (vtx,idx) counts come from descriptor tuples.

The +12 vertex offset is pinned by the recovered normals being exactly
unit-length (align16 lands 4 bytes early and silently corrupts every
field). A safety gate rejects any model whose indices are out of range
or whose mean |normal| is not ~1, declining garbage (Stage_S*
placeholders, the complex multi-stream hero-ship body) rather than
mis-decoding it.

- mesh.rs: Xbg7Model::from_xpr2 -> GameMesh { positions, normals, uvs,
  indices }; standalone f16->f32; unit + real-disc tests (weapon decodes
  to 215v/364t with unit normals + in-range UVs; DeltaSaber body
  declined).
- texture.rs: from_xpr2_index / texture_names so the viewer can pick a
  model's _col albedo map.
- viewer: loose .xpr with decodable XBG7 spawns Bevy meshes (real normals,
  double-sided) textured with the albedo, framed by the orbit camera; the
  central egui panel goes transparent so the 3D scene shows through.

Complex multi-stream body meshes (DeltaSaber f004, other vertex layouts)
remain undecoded and are cleanly declined — next target is dynamic RE.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-12 17:47:06 +02:00
MechaCat02
e6da726f7b docs(re): CONFIRM T8aD colour interpretation (linear UNORM, ARGB order)
Resolved the "colours look odd" channel-order/sRGB question for k_8_8_8_8
(T8aD 2D UI textures) with no emulator run:

- sRGB: Canary's Vulkan host-format table maps plain k_8_8_8_8 to
  R8G8B8A8_UNORM and contains no R8G8B8A8_SRGB anywhere — gamma is a
  separate explicit EDRAM path. So these are raw linear bytes; apply no
  sRGB/gamma decode.
- Channel order: measured 789 real disc textures — byte0 is the opaque
  alpha (==0xFF-dominant) in 97% — so on-disk is ARGB. Cross-checked
  against xenia-rs decode_k8888_tiled, the Canary-validated M1/M2 path,
  which nets the same [A,R,G,B]->[R,G,B,A]. sylpheed-formats is correct.

Promotes the k8888 entry HYPOTHESIS -> CONFIRMED; INDEX updated. XPR2
skybox tiling + viewer display-gamma remain separate open items.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-11 17:11:17 +02:00
33 changed files with 8394 additions and 187 deletions

40
Cargo.lock generated
View File

@@ -1892,6 +1892,25 @@ dependencies = [
"crossbeam-utils",
]
[[package]]
name = "crossbeam-deque"
version = "0.8.7"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "5181e0de7b61eb03a81e347d6dd8797bae9da5146707b51077e2d71a54ec0ceb"
dependencies = [
"crossbeam-epoch",
"crossbeam-utils",
]
[[package]]
name = "crossbeam-epoch"
version = "0.9.20"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "2d6914041f254d6e9176c01941b21115dcfb7089e55135a35411081bd106ef3f"
dependencies = [
"crossbeam-utils",
]
[[package]]
name = "crossbeam-utils"
version = "0.8.21"
@@ -4107,6 +4126,26 @@ version = "0.6.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "20675572f6f24e9e76ef639bc5552774ed45f1c30e2951e1e99c59888861c539"
[[package]]
name = "rayon"
version = "1.12.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "fb39b166781f92d482534ef4b4b1b2568f42613b53e5b6c160e24cfbfa30926d"
dependencies = [
"either",
"rayon-core",
]
[[package]]
name = "rayon-core"
version = "1.13.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "22e18b0f0062d30d4230b2e85ff77fdfe4326feb054b9783a3460d8435c8ab91"
dependencies = [
"crossbeam-deque",
"crossbeam-utils",
]
[[package]]
name = "redox_syscall"
version = "0.4.1"
@@ -4587,6 +4626,7 @@ dependencies = [
"binrw",
"flate2",
"futures",
"rayon",
"serde",
"serde_json",
"thiserror 2.0.18",

View File

@@ -97,6 +97,12 @@ enum Commands {
#[command(subcommand)]
cmd: PakCommands,
},
/// XBG7 mesh tools (inspect / headless render to PNG)
Mesh {
#[command(subcommand)]
cmd: MeshCommands,
},
}
#[derive(Subcommand)]
@@ -116,6 +122,51 @@ enum PakCommands {
/// Entry name-hash, e.g. `0x7c96296c`
hash: String,
},
/// Decode every T8aD texture in the pak (direct, RATC-nested, and LSTA
/// frames) to PNG — our decoder's output, for A/B against the running game.
Textures {
/// Path to the `.pak` index
pak: PathBuf,
/// Output directory for the PNGs (created if missing)
output: PathBuf,
/// Print per-texture dimensions + tile count.
#[arg(long)]
verbose: bool,
},
}
#[derive(Subcommand)]
enum MeshCommands {
/// Print the decoded sub-models of an XBG7 container (`.xpr`)
Info {
/// Path to the `.xpr` model / stage container
file: PathBuf,
},
/// Headless-render the decoded mesh(es) to a shaded PNG (software rasterizer)
Render {
/// Path to the `.xpr` model / stage container
file: PathBuf,
/// Output PNG path
output: PathBuf,
/// Image size in pixels (square)
#[arg(long, default_value_t = 900)]
size: u32,
/// Camera yaw in degrees
#[arg(long, default_value_t = 35.0)]
yaw: f32,
/// Camera pitch in degrees
#[arg(long, default_value_t = 22.0)]
pitch: f32,
/// Camera distance multiplier (1.0 = framed; <1 zooms in, >1 out)
#[arg(long, default_value_t = 1.0)]
dist: f32,
/// Force the stage grid layout even for single models
#[arg(long)]
row: bool,
/// Only render sub-models whose name contains this substring
#[arg(long)]
only: Option<String>,
},
}
#[derive(Subcommand)]
@@ -155,9 +206,18 @@ async fn main() -> Result<()> {
TextureCommands::Info { file } => cmd_texture_info(&file),
TextureCommands::Export { file, output } => cmd_texture_export(&file, &output),
},
Commands::Mesh { cmd } => match cmd {
MeshCommands::Info { file } => cmd_mesh_info(&file),
MeshCommands::Render { file, output, size, yaw, pitch, dist, row, only } => {
cmd_mesh_render(&file, &output, size, yaw, pitch, dist, row, only)
}
},
Commands::Pak { cmd } => match cmd {
PakCommands::List { pak, idxd_only } => cmd_pak_list(&pak, idxd_only),
PakCommands::Dump { pak, hash } => cmd_pak_dump(&pak, &hash),
PakCommands::List { pak, idxd_only } => cmd_pak_list(&pak, idxd_only),
PakCommands::Dump { pak, hash } => cmd_pak_dump(&pak, &hash),
PakCommands::Textures { pak, output, verbose } => {
cmd_pak_textures(&pak, &output, verbose)
}
},
}
}
@@ -357,6 +417,462 @@ fn cmd_texture_export(file: &Path, output: &Path) -> Result<()> {
Ok(())
}
// ── mesh info / render ──────────────────────────────────────────────────────
/// Decode a container and return its sub-models the same way the viewer routes:
/// a single model (weapon / prop) OR a stage's many sub-models — whichever
/// yields more geometry.
fn decode_models(bytes: &[u8]) -> Vec<sylpheed_formats::mesh::Xbg7Model> {
use sylpheed_formats::mesh::{count_xbg7, Xbg7Model};
// Route by container kind, NOT by whichever decoder yields more verts (that
// old heuristic let `stage_models`' content-anchoring win on single-model
// files, fabricating phantom / duplicate / mis-anchored blocks). A file with
// one XBG7 resource is a single model (weapon / prop) → use only the
// validated records-based list decode; content-anchoring a single-model file
// invents geometry. Many XBG7 resources → a Stage_* collection → anchor them.
if count_xbg7(bytes) > 1 {
// Multi-resource Stage_* collection → content-anchor every resource
// (each anchor now gated by stored-normal agreement).
Xbg7Model::stage_models(bytes)
} else if let Some(m) = Xbg7Model::from_xpr2(bytes)
.ok()
.filter(|m| !m.meshes.is_empty())
{
// Single model the records-based list decode carved (authoritative).
vec![m]
} else {
// Single model from_xpr2 couldn't locate (its sequential carve missed
// the block) — fall back to content-anchoring, which finds it by shape.
// Use a STRICT winding-consistency gate (0.85): on a single-model file a
// mis-anchor is an obvious phantom / spike-mess and must be declined,
// unlike the large stage corpus which keeps the ungated path.
Xbg7Model::anchor_models(bytes, 0.85)
}
}
fn cmd_mesh_info(file: &Path) -> Result<()> {
let bytes = std::fs::read(file).with_context(|| format!("Cannot read {}", file.display()))?;
let models = decode_models(&bytes);
if models.is_empty() {
println!("{} no decodable XBG7 geometry", "Mesh:".yellow().bold());
return Ok(());
}
let (mut tv, mut tt) = (0usize, 0usize);
println!("{} {}", "Mesh:".green().bold(), file.display());
for m in &models {
let (v, t) = m.totals();
tv += v;
tt += t;
let mut lo = [f32::MAX; 3];
let mut hi = [f32::MIN; 3];
for sub in &m.meshes {
for p in &sub.positions {
for a in 0..3 {
lo[a] = lo[a].min(p[a]);
hi[a] = hi[a].max(p[a]);
}
}
}
println!(
" {:16} {:>6} v {:>6} t bbox [{:.1} {:.1} {:.1}]",
m.name,
v,
t,
hi[0] - lo[0],
hi[1] - lo[1],
hi[2] - lo[2]
);
// Per-sub-mesh integrity diagnostics: degenerate triangles (a zero-area
// "hole"), vertices referenced by no triangle (dropped geometry), and the
// referenced index range vs the vertex count (short/over reads).
for (si, sub) in m.meshes.iter().enumerate() {
let nv = sub.positions.len();
let mut referenced = vec![false; nv];
let (mut degen, mut oob, mut imax) = (0usize, 0usize, 0u32);
for tri in sub.indices.chunks_exact(3) {
let (a, b, c) = (tri[0], tri[1], tri[2]);
imax = imax.max(a).max(b).max(c);
if a == b || b == c || a == c {
degen += 1;
}
for &i in tri {
if (i as usize) < nv {
referenced[i as usize] = true;
} else {
oob += 1;
}
}
}
let unref = referenced.iter().filter(|&&r| !r).count();
// Spanning triangles: longest edge ≫ the median (strip-junction spikes).
let edge = |a: u32, b: u32| {
let (p, q) = (sub.positions[a as usize], sub.positions[b as usize]);
((p[0] - q[0]).powi(2) + (p[1] - q[1]).powi(2) + (p[2] - q[2]).powi(2)).sqrt()
};
let mut maxedges: Vec<f32> = sub
.indices
.chunks_exact(3)
.map(|t| edge(t[0], t[1]).max(edge(t[1], t[2])).max(edge(t[0], t[2])))
.collect();
maxedges.sort_by(|a, b| a.partial_cmp(b).unwrap());
let median = maxedges.get(maxedges.len() / 2).copied().unwrap_or(1.0).max(1e-6);
let spanning = maxedges.iter().filter(|&&e| e > 6.0 * median).count();
println!(
" sub{si}: {nv} v, {} tris | degenerate {degen}, unref-verts {unref}, spanning>6×med {spanning}, idx_max {imax}/{}{}",
sub.indices.len() / 3,
nv.saturating_sub(1),
if oob > 0 { format!(", OOB {oob}") } else { String::new() },
);
// XDUMPVERT=1 → print the first few vertex positions per sub-mesh, for
// content-matching a decoded sub-mesh against the GPU draw log.
if std::env::var("XDUMPVERT").is_ok() {
let mut lo = [f32::MAX; 3];
let mut hi = [f32::MIN; 3];
for p in &sub.positions {
for a in 0..3 {
lo[a] = lo[a].min(p[a]);
hi[a] = hi[a].max(p[a]);
}
}
println!(
" bbox X[{:.2}..{:.2}] Y[{:.2}..{:.2}] Z[{:.2}..{:.2}] ctr({:.2},{:.2},{:.2})",
lo[0], hi[0], lo[1], hi[1], lo[2], hi[2],
(lo[0]+hi[0])/2.0, (lo[1]+hi[1])/2.0, (lo[2]+hi[2])/2.0
);
}
}
}
println!(
" {} {} sub-models · {} verts · {} tris",
"TOTAL".bold(),
models.len(),
tv,
tt
);
Ok(())
}
#[allow(clippy::too_many_arguments)]
fn cmd_mesh_render(
file: &Path,
output: &Path,
size: u32,
yaw: f32,
pitch: f32,
dist: f32,
force_row: bool,
only: Option<String>,
) -> Result<()> {
let bytes = std::fs::read(file).with_context(|| format!("Cannot read {}", file.display()))?;
let mut models = decode_models(&bytes);
if let Some(sub) = &only {
// Prefer an exact name match (e.g. `f001` for the neutral ship pose,
// excluding the `_rou_f001_mnv*` animation poses that also *contain*
// "f001"); fall back to substring when nothing matches exactly.
if models.iter().any(|m| m.name == *sub) {
models.retain(|m| m.name == *sub);
} else {
models.retain(|m| m.name.contains(sub.as_str()));
}
}
if models.is_empty() {
anyhow::bail!("no decodable XBG7 geometry in {}", file.display());
}
// ── Build a triangle soup. ──
// Single models render centred; multi-model containers (stages) get the
// viewer's normalised **thumbnail grid**: each sub-model recentred and
// uniformly scaled to a fixed cell, so all are equally visible regardless of
// native scale (mirrors `spawn_stage_models`).
let multi = models.len() > 1 || force_row;
// XMIRROR=x|y|z → negate that axis, to test an Xbox(LH)→Bevy(RH) handedness
// flip against reference screenshots.
let mirror: [f32; 3] = match std::env::var("XMIRROR").ok().as_deref() {
Some("x") => [-1.0, 1.0, 1.0],
Some("y") => [1.0, -1.0, 1.0],
Some("z") => [1.0, 1.0, -1.0],
_ => [1.0, 1.0, 1.0],
};
let mut tris: Vec<[[f32; 3]; 3]> = Vec::new();
// XCOLORSUB=1 tints each sub-mesh a distinct colour (to see which sub is
// which part / where the "extra fin" comes from). Parallel to `tris`.
let color_sub = std::env::var("XCOLORSUB").is_ok();
// XONLYSUB=N renders only the N-th global sub-mesh (to isolate one part).
let only_sub: Option<usize> = std::env::var("XONLYSUB").ok().and_then(|s| s.parse().ok());
let mut tints: Vec<[f32; 3]> = Vec::new();
const PALETTE: [[f32; 3]; 8] = [
[1.0, 1.0, 1.0], // sub0 body = white
[1.0, 0.35, 0.35], // sub1 red
[0.35, 1.0, 0.35], // sub2 green
[0.4, 0.55, 1.0], // sub3 blue
[1.0, 0.9, 0.3], // sub4 yellow
[1.0, 0.5, 1.0], // sub5 magenta
[0.3, 1.0, 1.0], // sub6 cyan
[1.0, 0.6, 0.2], // sub7 orange
];
let mut sub_gi = 0usize;
const CELL: f32 = 10.0;
const GAP: f32 = 4.0;
let grid_pitch = CELL + GAP;
let cols = (models.len() as f32).sqrt().ceil().max(1.0) as usize;
for (i, m) in models.iter().enumerate() {
let mut lo = [f32::MAX; 3];
let mut hi = [f32::MIN; 3];
for sub in &m.meshes {
for p in &sub.positions {
for a in 0..3 {
lo[a] = lo[a].min(p[a]);
hi[a] = hi[a].max(p[a]);
}
}
}
if lo[0] > hi[0] {
continue;
}
let center = [
(lo[0] + hi[0]) * 0.5,
(lo[1] + hi[1]) * 0.5,
(lo[2] + hi[2]) * 0.5,
];
let (scale, cell) = if multi {
let extent = (hi[0] - lo[0]).max(hi[1] - lo[1]).max(hi[2] - lo[2]).max(1e-3);
let col = i % cols;
let row = i / cols;
(CELL / extent, [col as f32 * grid_pitch, -(row as f32) * grid_pitch, 0.0])
} else {
(1.0, [0.0, 0.0, 0.0])
};
// XNODEXFORM=1 applies the XBG7 scene-graph node placement (fins move to
// the tail) — to verify the transforms recovered from the graph.
let placements = if std::env::var("XNODEXFORM").is_ok() {
sylpheed_formats::mesh::node_transforms(&bytes, &m.name)
} else {
Vec::new()
};
for (sub_local, sub) in m.meshes.iter().enumerate() {
// Every scene-graph instance that draws this sub-mesh (mirrored fin
// pair, L/R winglets…); `None` = no graph placement → identity.
let mine: Vec<Option<&sylpheed_formats::mesh::NodePlacement>> = {
let v: Vec<_> = placements
.iter()
.filter(|p| p.sub_index == sub_local)
.map(Some)
.collect();
if v.is_empty() {
vec![None]
} else {
v
}
};
// Sub-mesh indices are a triangle list (the decoder has already
// expanded the file's triangle strips).
let n = sub.positions.len();
// XSPANONLY=1 renders ONLY long-edge ("spanning") triangles; XSPANHIDE=1
// renders everything EXCEPT them — to see whether the flagged spanning
// triangles are real geometry or decode artifacts (phantom sheets).
let span_only = std::env::var("XSPANONLY").is_ok();
let span_hide = std::env::var("XSPANHIDE").is_ok();
let med = {
let mut e: Vec<f32> = sub
.indices
.chunks_exact(3)
.filter(|t| (t[0] as usize) < n && (t[1] as usize) < n && (t[2] as usize) < n)
.map(|t| {
let d = |a: u32, b: u32| {
let (p, q) = (sub.positions[a as usize], sub.positions[b as usize]);
((p[0] - q[0]).powi(2) + (p[1] - q[1]).powi(2) + (p[2] - q[2]).powi(2))
.sqrt()
};
d(t[0], t[1]).max(d(t[1], t[2])).max(d(t[0], t[2]))
})
.collect();
e.sort_by(|a, b| a.partial_cmp(b).unwrap());
e.get(e.len() / 2).copied().unwrap_or(1.0).max(1e-6)
};
if let Some(want) = only_sub {
if sub_gi != want {
sub_gi += 1;
continue;
}
}
let tint = if color_sub {
PALETTE[sub_gi % PALETTE.len()]
} else {
[1.0, 1.0, 1.0]
};
for place in &mine {
let f = |i: usize| {
let p = place.map(|pl| pl.apply(sub.positions[i])).unwrap_or(sub.positions[i]);
[
(p[0] - center[0]) * scale * mirror[0] + cell[0],
(p[1] - center[1]) * scale * mirror[1] + cell[1],
(p[2] - center[2]) * scale * mirror[2] + cell[2],
]
};
for tri in sub.indices.chunks_exact(3) {
let (a, b, c) = (tri[0] as usize, tri[1] as usize, tri[2] as usize);
if a < n && b < n && c < n {
if span_only || span_hide {
let d = |i: usize, j: usize| {
let (p, q) = (sub.positions[i], sub.positions[j]);
((p[0] - q[0]).powi(2)
+ (p[1] - q[1]).powi(2)
+ (p[2] - q[2]).powi(2))
.sqrt()
};
let spanning = d(a, b).max(d(b, c)).max(d(a, c)) > 6.0 * med;
if span_only && !spanning {
continue;
}
if span_hide && spanning {
continue;
}
}
tris.push([f(a), f(b), f(c)]);
tints.push(tint);
}
}
}
sub_gi += 1;
}
}
if tris.is_empty() {
anyhow::bail!("no triangles to render");
}
let rgba = rasterize(&tris, &tints, size, yaw, pitch, dist);
image::save_buffer(output, &rgba, size, size, image::ExtendedColorType::Rgba8)
.with_context(|| format!("writing PNG {}", output.display()))?;
println!(
"{} {} tris → {} ({}×{}, yaw {:.0}° pitch {:.0}°)",
"Rendered".green().bold(),
tris.len(),
output.display().to_string().cyan(),
size,
size,
yaw,
pitch,
);
Ok(())
}
/// Minimal software rasterizer: orthographic, z-buffered, two-sided Lambert +
/// headlight shading over a flat grey material on a dark background. Enough to
/// judge whether recovered geometry is coherent.
fn rasterize(
tris: &[[[f32; 3]; 3]],
tints: &[[f32; 3]],
size: u32,
yaw_deg: f32,
pitch_deg: f32,
dist: f32,
) -> Vec<u8> {
let n = size as usize;
let (yaw, pitch) = (yaw_deg.to_radians(), pitch_deg.to_radians());
let (cy, sy) = (yaw.cos(), yaw.sin());
let (cp, sp) = (pitch.cos(), pitch.sin());
// Rotate a world point into view space (yaw about Y, then pitch about X).
let view = |p: [f32; 3]| -> [f32; 3] {
let x = p[0] * cy + p[2] * sy;
let z0 = -p[0] * sy + p[2] * cy;
let y = p[1] * cp - z0 * sp;
let z = p[1] * sp + z0 * cp;
[x, y, z]
};
// View-space bbox → orthographic fit.
let mut lo = [f32::MAX; 3];
let mut hi = [f32::MIN; 3];
for t in tris {
for v in t {
let q = view(*v);
for a in 0..3 {
lo[a] = lo[a].min(q[a]);
hi[a] = hi[a].max(q[a]);
}
}
}
let span = (hi[0] - lo[0]).max(hi[1] - lo[1]).max(1e-3);
let scale = (n as f32) * 0.9 / (span * dist.max(1e-3));
let cx = (lo[0] + hi[0]) * 0.5;
let cyv = (lo[1] + hi[1]) * 0.5;
let to_screen = |q: [f32; 3]| -> (f32, f32, f32) {
let sx = (q[0] - cx) * scale + n as f32 * 0.5;
let sy = n as f32 * 0.5 - (q[1] - cyv) * scale;
(sx, sy, q[2])
};
let mut color = vec![18u8; n * n * 4];
for i in 0..n * n {
color[i * 4 + 3] = 255;
}
let mut depth = vec![f32::MAX; n * n];
// Light in view space (upper-left-front).
let light = {
let l = [-0.4f32, 0.6, 0.7];
let m = (l[0] * l[0] + l[1] * l[1] + l[2] * l[2]).sqrt();
[l[0] / m, l[1] / m, l[2] / m]
};
for (ti, t) in tris.iter().enumerate() {
let tint = tints.get(ti).copied().unwrap_or([1.0, 1.0, 1.0]);
let v0 = view(t[0]);
let v1 = view(t[1]);
let v2 = view(t[2]);
// Face normal in view space.
let e1 = [v1[0] - v0[0], v1[1] - v0[1], v1[2] - v0[2]];
let e2 = [v2[0] - v0[0], v2[1] - v0[1], v2[2] - v0[2]];
let mut nrm = [
e1[1] * e2[2] - e1[2] * e2[1],
e1[2] * e2[0] - e1[0] * e2[2],
e1[0] * e2[1] - e1[1] * e2[0],
];
let nl = (nrm[0] * nrm[0] + nrm[1] * nrm[1] + nrm[2] * nrm[2]).sqrt();
if nl < 1e-12 {
continue;
}
nrm = [nrm[0] / nl, nrm[1] / nl, nrm[2] / nl];
// Two-sided: diffuse from |n·L|, plus a headlight term from |n.z|.
let diff = (nrm[0] * light[0] + nrm[1] * light[1] + nrm[2] * light[2]).abs();
let head = nrm[2].abs();
let inten = (0.18 + 0.55 * diff + 0.3 * head).min(1.0);
let shade = (inten * 210.0) as u8;
let (ax, ay, az) = to_screen(v0);
let (bx, by, bz) = to_screen(v1);
let (ccx, ccy, ccz) = to_screen(v2);
let minx = ax.min(bx).min(ccx).floor().max(0.0) as usize;
let maxx = ax.max(bx).max(ccx).ceil().min(n as f32 - 1.0) as usize;
let miny = ay.min(by).min(ccy).floor().max(0.0) as usize;
let maxy = ay.max(by).max(ccy).ceil().min(n as f32 - 1.0) as usize;
let area = (bx - ax) * (ccy - ay) - (by - ay) * (ccx - ax);
if area.abs() < 1e-6 {
continue;
}
for py in miny..=maxy {
for px in minx..=maxx {
let fx = px as f32 + 0.5;
let fy = py as f32 + 0.5;
let w0 = ((bx - fx) * (ccy - fy) - (by - fy) * (ccx - fx)) / area;
let w1 = ((ccx - fx) * (ay - fy) - (ccy - fy) * (ax - fx)) / area;
let w2 = 1.0 - w0 - w1;
if w0 < 0.0 || w1 < 0.0 || w2 < 0.0 {
continue;
}
let z = w0 * az + w1 * bz + w2 * ccz;
let idx = py * n + px;
if z < depth[idx] {
depth[idx] = z;
color[idx * 4] = (shade as f32 * tint[0]).min(255.0) as u8;
color[idx * 4 + 1] = (shade as f32 * tint[1]).min(255.0) as u8;
color[idx * 4 + 2] = (shade as f32 * tint[2] * 1.02).min(255.0) as u8;
}
}
}
}
color
}
/// Software-decode a de-tiled `X360Texture` (mip 0) to tightly-packed RGBA8.
///
/// BCn blocks are decompressed with `texpresso`; uncompressed A8R8G8B8/X8R8G8B8
@@ -523,3 +1039,179 @@ fn cmd_pak_dump(pak: &Path, hash_str: &str) -> Result<()> {
}
Ok(())
}
// ── pak textures ─────────────────────────────────────────────────────────────
/// Turn a child name into a filesystem-safe fragment.
fn safe_name(s: &str) -> String {
s.chars()
.map(|c| {
if c.is_ascii_alphanumeric() || matches!(c, '.' | '_' | '-') {
c
} else {
'_'
}
})
.collect()
}
#[derive(Default)]
struct TexStats {
written: usize,
skipped: usize,
}
fn be32_at(b: &[u8], off: usize) -> u32 {
u32::from_be_bytes([b[off], b[off + 1], b[off + 2], b[off + 3]])
}
/// Decode one T8aD slice (whose first bytes are the magic) to PNG. With
/// `verbose`, prints its dimensions + tile count; the `XDUMPHDR` env var dumps
/// the raw header (base + offset table) for format RE.
fn emit_t8ad(
slice: &[u8],
hash: u32,
stem: &str,
output: &Path,
verbose: bool,
stats: &mut TexStats,
) -> Result<()> {
use sylpheed_formats::t8ad;
if !t8ad::is_t8ad(slice) || slice.len() < 0x40 {
return Ok(());
}
let (w, h, tiles) = (
be32_at(slice, 0x14),
be32_at(slice, 0x18),
be32_at(slice, 0x1c),
);
// Debug: dump the header — 44-byte base + the `tiles`-entry u32 offset table.
if std::env::var("XDUMPHDR").is_ok() {
println!("\n{stem} {w}x{h} tiles={tiles}");
print!(" base[0x00..0x2c]:");
for i in (0..44).step_by(4) {
print!(" {:08x}", be32_at(slice, i));
}
print!("\n offsets:");
for t in 0..(tiles as usize).min(64) {
if 0x2c + t * 4 + 4 <= slice.len() {
print!(" {}", be32_at(slice, 0x2c + t * 4));
}
}
println!();
return Ok(());
}
if verbose {
println!(" {hash:08x} {stem:<34} {w:>4}x{h:<4} tiles {tiles}");
}
match t8ad::parse(slice) {
Some(img) => {
let out =
output.join(format!("{hash:08x}_{stem}_{}x{}.png", img.width, img.height));
image::save_buffer(
&out,
&img.rgba,
img.width,
img.height,
image::ExtendedColorType::Rgba8,
)
.with_context(|| format!("writing PNG {}", out.display()))?;
stats.written += 1;
}
None => stats.skipped += 1,
}
Ok(())
}
/// Decode every T8aD in a pak (direct entries, RATC-nested children, LSTA
/// frames) to PNG — our decoder's exact output — for A/B against the game.
fn cmd_pak_textures(pak: &Path, output: &Path, verbose: bool) -> Result<()> {
use sylpheed_formats::{lsta, ratc, t8ad};
let arc = PakArchive::open(pak).with_context(|| format!("opening {}", pak.display()))?;
std::fs::create_dir_all(output)
.with_context(|| format!("creating {}", output.display()))?;
println!(
"{} {}{}",
"Textures".green().bold(),
pak.display().to_string().cyan(),
output.display().to_string().cyan(),
);
let mut stats = TexStats::default();
for e in arc.entries() {
let payload = match arc.read(e) {
Ok(p) => p,
Err(_) => continue,
};
let hash = e.name_hash;
// Direct T8aD entry.
if t8ad::is_t8ad(&payload) {
emit_t8ad(&payload, hash, "direct", output, verbose, &mut stats)?;
continue;
}
// LSTA sprite list = N inline T8aD frames (walk by magic, emit each).
if lsta::is_lsta(&payload) {
let mut off = 0usize;
let mut idx = 0usize;
while let Some(pos) = payload[off..]
.windows(4)
.position(|w| w == &t8ad::T8AD_MAGIC)
{
let start = off + pos;
let next = payload[start + 4..]
.windows(4)
.position(|w| w == &t8ad::T8AD_MAGIC)
.map(|p| start + 4 + p)
.unwrap_or(payload.len());
emit_t8ad(
&payload[start..next],
hash,
&format!("lsta{idx:03}"),
output,
verbose,
&mut stats,
)?;
idx += 1;
off = next;
}
continue;
}
// RATC bundle: decode its T8aD children (named, e.g. `foo.t32`).
if ratc::is_ratc(&payload) {
if let Some(children) = ratc::parse(&payload) {
for (i, child) in children.iter().enumerate() {
if child.kind != "T8aD" {
continue;
}
let end = (child.offset + child.size).min(payload.len());
if child.offset >= end {
continue;
}
let stem = if child.name.is_empty() {
format!("child{i:03}")
} else {
safe_name(&child.name)
};
emit_t8ad(&payload[child.offset..end], hash, &stem, output, verbose, &mut stats)?;
}
}
continue;
}
}
println!(
"\n {} PNG(s) written, {} undecodable (non-tilecount variants — likely DXT)",
stats.written.to_string().green(),
stats.skipped.to_string().yellow(),
);
Ok(())
}

View File

@@ -19,5 +19,11 @@ tracing = { workspace = true }
serde = { workspace = true }
serde_json = { workspace = true }
# Native-only: parallelise the per-resource XBG7 stage anchoring (hundreds of
# independent sub-models per container). rayon needs OS threads, so it is not
# pulled in for the wasm32 build, which falls back to a sequential decode.
[target.'cfg(not(target_arch = "wasm32"))'.dependencies]
rayon = "1"
[dev-dependencies]
tokio = { workspace = true }

View File

@@ -0,0 +1,36 @@
use sylpheed_formats::{hash::name_hash, movie_manifest, movie_voice, slb, PakArchive};
use std::fs;use std::io::{Read,Seek,SeekFrom};use std::process::Command;
fn rg(disc:&str,g:u64,n:usize)->Vec<u8>{let mut segs=vec![];let mut cum=0u64;for i in 0..5{let p=format!("{disc}/dat/sound.p{i:02}");if let Ok(m)=fs::metadata(&p){segs.push((cum,m.len(),p));cum+=m.len();}}let mut out=vec![];let(mut need,mut pos)=(n,g);for(base,len,path)in &segs{if need==0||pos>=base+len||pos<*base{continue;}let local=pos-base;let take=need.min((len-local)as usize);let mut f=fs::File::open(path).unwrap();f.seek(SeekFrom::Start(local)).unwrap();let mut b=vec![0u8;take];f.read_exact(&mut b).unwrap();out.extend_from_slice(&b);need-=take;pos+=take as u64;}out}
fn contains(h:&[u8],n:&[u8])->bool{h.windows(n.len()).any(|w|w==n)}
fn dur(w:&str)->String{let o=Command::new("ffprobe").args(["-v","error","-show_entries","format=duration","-of","default=nw=1:nk=1",w]).output().unwrap();String::from_utf8_lossy(&o.stdout).trim().to_string()}
fn hokyu_fallback(m:&str)->Option<String>{if !m.starts_with("hokyu_"){return None}let ls=m.contains("_LS_");let ds=m.contains("_DS_");let a=m.ends_with('A');let h=m.ends_with('H');Some(match(ls,ds,a,h){(true,_,true,_)=>"VOICE_D_450",(true,_,_,true)=>"VOICE_D_453",(_,true,true,_)=>"VOICE_D_452",(_,true,_,true)=>"VOICE_D_454",_=>return None}.into())}
fn main(){
let disc=std::env::var("SYLPHEED_DISC").unwrap();
let tpak=PakArchive::open(format!("{disc}/dat/tables.pak")).unwrap();
let manifest=tpak.entries().iter().find_map(|e|tpak.read(e).ok().filter(|b|movie_manifest::is_manifest(b))).unwrap();
let code="eng";
let registry=tpak.entries().iter().find_map(|e|tpak.read(e).ok().filter(|b|contains(b,format!("{code}\\Movie\\VOICE_ADV.slb").as_bytes()))).unwrap();
let ids=movie_voice::registry_voice_ids(&registry);
let stoc=fs::read(format!("{disc}/dat/sound.pak")).unwrap();
let entries=PakArchive::parse_toc(&stoc).unwrap();
let hokyu:Vec<String>=movie_manifest::parse(&manifest).into_iter().filter(|m|m.movie.starts_with("hokyu_")).map(|m|m.movie).collect();
for movie in &hokyu{
let bound=movie_manifest::voice_token(&manifest,movie);
let token=bound.clone().or_else(||hokyu_fallback(movie));
let Some(token)=token else{println!("{movie:16} NO TOKEN"); continue};
let src = if bound.is_some(){"bound"}else{"fallback"};
let Some(&id)=ids.get(&token) else{println!("{movie:16} {token} not in registry");continue};
let Some(anchor)=["Movie","etc","Voice"].iter().find_map(|d|{let h=name_hash(&format!("{code}\\{d}\\{token}.slb"));entries.binary_search_by_key(&h,|e|e.name_hash).ok().map(|i|entries[i].offset as u64)}) else{println!("{movie:16} no anchor");continue};
let win_start=(anchor.saturating_sub(2*1024*1024))&!3;
let window=rg(&disc,win_start,8*1024*1024);
let Some(el)=movie_voice::find_descriptor(&window,id) else{println!("{movie:16} desc not found");continue};
let end=win_start+el as u64;
let start=movie_voice::find_descriptor(&window,id.wrapping_sub(1)).or_else(||movie_voice::find_descriptor_before(&window,el)).map(|o|win_start+o as u64).filter(|&s|s<end&&end-s<1_500_000).unwrap_or(anchor);
let region=rg(&disc,start,(end-start)as usize);
let mut riffs=slb::to_xma_riffs(&region); if riffs.is_empty(){riffs=slb::to_xma_riff_best(&region).into_iter().collect();}
let mut d="?".into();
if let Some(r)=riffs.first(){let xp=format!("/tmp/hf_{movie}.xma.wav");let wp=format!("/tmp/hf_{movie}.wav");fs::write(&xp,r).unwrap();let _=Command::new("ffmpeg").args(["-hide_banner","-v","error","-y","-i",&xp,&wp]).status();d=dur(&wp);}
let mv=dur(&format!("{disc}/dat/movie/{movie}.wmv"));
println!("{movie:16} {src:8} {token:12} id={id} voice={d}s movie={mv}s");
}
}

View File

@@ -0,0 +1,46 @@
//! Validation: dump the parsed mission/phase cutscene map from the real manifest.
use sylpheed_formats::movie_manifest::{self, MovieKind};
use sylpheed_formats::pak::PakArchive;
fn main() {
let disc = std::env::var("SYLPHEED_DISC").expect("set SYLPHEED_DISC");
let arc = PakArchive::open(format!("{disc}/dat/tables.pak")).unwrap();
let man = arc
.entries()
.iter()
.find_map(|e| arc.read(e).ok().filter(|b| movie_manifest::is_manifest(b)))
.expect("manifest");
let entries = movie_manifest::parse(&man);
let mut counts = [0usize; 5];
for e in &entries {
counts[match e.kind {
MovieKind::System => 0,
MovieKind::Intro => 1,
MovieKind::Phase => 2,
MovieKind::PhaseEnd => 3,
MovieKind::Supply => 4,
}] += 1;
let m = e.mission.map(|x| x.to_string()).unwrap_or_default();
let p = e.phase.map(|x| x.to_string()).unwrap_or_default();
println!(
"{:<22} {:<9} m={:<2} p={:<1} {:<16} {:<14} tel={:<14} sub={}",
e.slot,
format!("{:?}", e.kind),
m,
p,
e.movie,
e.voice_token.as_deref().unwrap_or("-"),
e.telop.as_deref().unwrap_or("-"),
e.subtitle.as_deref().unwrap_or("-"),
);
}
eprintln!(
"TOTAL {} entries — System={} Intro={} Phase={} PhaseEnd={} Supply={}",
entries.len(),
counts[0],
counts[1],
counts[2],
counts[3],
counts[4]
);
}

View File

@@ -0,0 +1,36 @@
use sylpheed_formats::{hash::name_hash, movie_manifest, movie_voice, slb, PakArchive};
use std::fs;use std::io::{Read,Seek,SeekFrom};use std::process::Command;
fn rg(disc:&str,g:u64,n:usize)->Vec<u8>{let mut segs=vec![];let mut cum=0u64;for i in 0..5{let p=format!("{disc}/dat/sound.p{i:02}");if let Ok(m)=fs::metadata(&p){segs.push((cum,m.len(),p));cum+=m.len();}}let mut out=vec![];let(mut need,mut pos)=(n,g);for(base,len,path)in &segs{if need==0||pos>=base+len||pos<*base{continue;}let local=pos-base;let take=need.min((len-local)as usize);let mut f=fs::File::open(path).unwrap();f.seek(SeekFrom::Start(local)).unwrap();let mut b=vec![0u8;take];f.read_exact(&mut b).unwrap();out.extend_from_slice(&b);need-=take;pos+=take as u64;}out}
fn contains(h:&[u8],n:&[u8])->bool{h.windows(n.len()).any(|w|w==n)}
fn dur(w:&str)->String{let o=Command::new("ffprobe").args(["-v","error","-show_entries","format=duration","-of","default=nw=1:nk=1",w]).output().unwrap();String::from_utf8_lossy(&o.stdout).trim().to_string()}
fn main(){
let disc=std::env::var("SYLPHEED_DISC").unwrap();
let tpak=PakArchive::open(format!("{disc}/dat/tables.pak")).unwrap();
let manifest=tpak.entries().iter().find_map(|e|tpak.read(e).ok().filter(|b|movie_manifest::is_manifest(b))).unwrap();
let code="eng";
let registry=tpak.entries().iter().find_map(|e|tpak.read(e).ok().filter(|b|contains(b,format!("{code}\\Movie\\VOICE_ADV.slb").as_bytes()))).unwrap();
let ids=movie_voice::registry_voice_ids(&registry);
let stoc=fs::read(format!("{disc}/dat/sound.pak")).unwrap();
let entries=PakArchive::parse_toc(&stoc).unwrap();
for movie in ["ADV","RT01A","RT01B","RT01C_1","S00A","S01A","hokyu_LS_s02A","hokyu_LS_s09A","hokyu_LS_s02H","hokyu_DS_s07H"]{
let Some(token)=movie_manifest::voice_token(&manifest,movie) else { println!("{movie:16} no token"); continue };
let Some(&id)=ids.get(&token) else { println!("{movie:16} token {token} not in registry"); continue };
let Some(anchor)=["Movie","etc","Voice"].iter().find_map(|d|{let h=name_hash(&format!("{code}\\{d}\\{token}.slb"));entries.binary_search_by_key(&h,|e|e.name_hash).ok().map(|i|entries[i].offset as u64)}) else { println!("{movie:16} no anchor for {token}"); continue };
let win_start=(anchor.saturating_sub(2*1024*1024))&!3;
let window=rg(&disc,win_start,8*1024*1024);
let Some(end_local)=movie_voice::find_descriptor(&window,id) else { println!("{movie:16} id {id}: desc NOT FOUND"); continue };
let end=win_start+end_local as u64;
let start=movie_voice::find_descriptor(&window,id.wrapping_sub(1))
.or_else(||movie_voice::find_descriptor_before(&window,end_local))
.map(|o|win_start+o as u64)
.filter(|&s|s<end && end-s<1_500_000)
.unwrap_or(anchor);
let region=rg(&disc,start,(end-start) as usize);
let mut riffs=slb::to_xma_riffs(&region);
if riffs.is_empty(){ riffs=slb::to_xma_riff_best(&region).into_iter().collect(); }
let mut d="?".into();
if let Some(r)=riffs.first(){ let xp=format!("/tmp/ra2_{movie}.xma.wav");let wp=format!("/tmp/ra2_{movie}.wav");fs::write(&xp,r).unwrap();let _=Command::new("ffmpeg").args(["-hide_banner","-v","error","-y","-i",&xp,&wp]).status();d=dur(&wp);}
let mv=dur(&format!("{disc}/dat/movie/{movie}.wmv"));
println!("{movie:16} id={id:5} region={}KB riffs={} voice={d}s movie={mv}s", (end-start)/1024, riffs.len());
}
}

View File

@@ -0,0 +1,32 @@
use sylpheed_formats::slb;
use std::fs;use std::io::{Read,Seek,SeekFrom};use std::process::Command;
fn rg(disc:&str,g:u64,n:usize)->Vec<u8>{let mut segs=vec![];let mut cum=0u64;for i in 0..5{let p=format!("{disc}/dat/sound.p{i:02}");if let Ok(m)=fs::metadata(&p){segs.push((cum,m.len(),p));cum+=m.len();}}let mut out=vec![];let(mut need,mut pos)=(n,g);for(base,len,path)in &segs{if need==0||pos>=base+len||pos<*base{continue;}let local=pos-base;let take=need.min((len-local)as usize);let mut f=fs::File::open(path).unwrap();f.seek(SeekFrom::Start(local)).unwrap();let mut b=vec![0u8;take];f.read_exact(&mut b).unwrap();out.extend_from_slice(&b);need-=take;pos+=take as u64;}out}
fn dur(w:&str)->String{let o=Command::new("ffprobe").args(["-v","error","-show_entries","format=duration:stream=channels","-of","default=nw=1:nk=1",w]).output().unwrap();String::from_utf8_lossy(&o.stdout).replace('\n'," ")}
fn movdur(disc:&str,m:&str)->String{ dur(&format!("{disc}/dat/movie/{m}")) }
fn main(){
let disc=std::env::var("SYLPHEED_DISC").unwrap();
// descriptor trailer global offsets (from scan) => cue N data = [desc(N-1)..desc(N)]
// (id, end_off, movie)
let cues=[(1600u32,437044592u64,"ADV.wmv"),(1601,437345648,"RT01A.wmv"),(1602,437712240,"RT01B.wmv"),
(1603,438080880,"RT01C_1.wmv"),(1604,438451568,"RT01C_2.wmv"),(1605,438789488,"RT02A.wmv")];
for k in 1..cues.len(){
let (id,end,mov)=cues[k];
let start=cues[k-1].1;
let region=rg(&disc,start,(end-start) as usize + 12000); // +tail to include full data before next trailer
let riffs=slb::to_xma_riffs(&region);
let mut total=0.0f32; let mut parts=vec![];
for (j,r) in riffs.iter().enumerate(){
let xp=format!("/tmp/cue{id}_{j}.xma.wav"); let wp=format!("/tmp/cue{id}_{j}.wav");
fs::write(&xp,r).unwrap();
let _=Command::new("ffmpeg").args(["-hide_banner","-v","error","-y","-i",&xp,&wp]).status();
let d=dur(&wp); parts.push(d.clone());
total+=d.split_whitespace().next().unwrap_or("0").parse::<f32>().unwrap_or(0.0);
}
// spanning?
let span_adv_end=437547264u64;
let spans = start < span_adv_end && end > span_adv_end || (start/1_000 != end/1_000);
println!("cue {id} ({mov}): region[{start}..{end}] {} bytes, {} riff(s), Σ={:.1}s | movie={} parts={:?}",
end-start, riffs.len(), total, movdur(&disc,mov), parts);
}
println!("\nWAVs at /tmp/cue16XX_*.wav (listen: RT01B=1602, RT01C_1=1603, RT01C_2=1604)");
}

View File

@@ -50,10 +50,23 @@ pub mod mesh;
// Audio parsing — scaffold for XMA handling
pub mod audio;
// Movie cutscene subtitles — the movie → track → text chain.
pub mod movie_subtitle;
// XACT `.slb` sound banks (voice / music) → decodable XMA1 RIFF.
pub mod slb;
// The movie manifest: authoritative movie → subtitle → voice index.
pub mod movie_manifest;
pub mod movie_voice;
/// Re-export the most commonly used types at the crate root.
pub use font::FontInfo;
pub use idxd::{IdxdError, IdxdObject};
pub use ixud::{Cue, Subtitle};
pub use movie_subtitle::{SubCue, SubLang};
pub use mesh::{GameMesh, Xbg7Model};
pub use ratc::RatcChild;
pub use t8ad::T8adImage;
pub use pak::{PakArchive, PakEntry, PakError};

File diff suppressed because it is too large Load Diff

View File

@@ -0,0 +1,428 @@
//! The movie manifest: the game's authoritative **mission → phase → cutscene**
//! table (statically reversed). Beyond the movie→voice binding, it defines the
//! whole cutscene *structure*: which video (+ subtitle, + on-screen text overlay,
//! + voice) plays at each mission phase, and of what kind (story intro, in-mission
//! phase cutscene, resupply, …).
//!
//! ## On-disc structure (`dat/tables.pak`, entry hash `0x5b983a08`)
//!
//! One `IDXD` table (`Z1`+zlib block), schema `0x067025b9`. Unlike a flat IDXD
//! object, this one is an **array of records**, and its string pool is laid out
//! as **two consecutive arrays**:
//! 1. the **slot KEYS** — `LOGO1`, `ADVERTISE_MOVIE`, `STAFF_ROLL`, then one
//! per cutscene: `MS<NN><X>`, `STAGE<NN>_PHASE0<N>`, `STAGE<NN>_PHASE_END…`,
//! `S<NN>_SUPPLY_<ACROPOLIS|TANKER>`, in play order;
//! 2. the **value RECORDS**, in the same order — each a run led by `<movie>.wmv`
//! and (optionally) a `<pak>+…​.prt` overlay (field `TELOP`), a
//! `<pak>+SUBTITLE_<movie>.tbl` (field `SUBTITLE`), and a `VOICE_<token>`
//! (field `VOICETRACK`). A few slot keys are **valueless** (e.g. stages 5 and
//! 12 have no resupply video) — those slots have no record and are skipped.
//!
//! The **slot key is the semantic role** — this is how the game separates a
//! cutscene voice from in-mission radio chatter (structurally, by which slot/bank,
//! not one boolean): `MS*` = story intro (`VOICE_S*`), `STAGE*_PHASE*` = radio-box
//! cutscene (`VOICE_RT*`), `S*_SUPPLY_*` = resupply (`VOICE_D_45x` or none).
//! Resupply videos are **reused across stages** (`S04_SUPPLY_ACROPOLIS` →
//! `hokyu_DS_s02A`), so the mapping is genuinely a table, not derivable from names.
//!
//! The voice token is likewise **not** always `VOICE_<movie>`: most movies use
//! `VOICE_<movie>` (in `<lang>\Movie\`), but the resupply cutscenes bind to
//! in-mission radio clips like `VOICE_D_450` (in `<lang>\etc\`), and some have no
//! voice at all. This manifest is the ground truth.
//!
//! ## Alignment
//!
//! Slot keys and value records would be a 1:1 positional zip but for the valueless
//! slots. We recover the gaps without decoding the binary node/index region by
//! using the fully-derivable **`MS<NN><X>` → `S<NN><X>` anchors**: a non-`MS` slot
//! whose record would be an `MS` target actually belongs to the upcoming `MS`
//! slot, so the non-`MS` slot is a gap. (Decoding the node region would make this
//! exact and also unlock the defaulted numeric fields in other IDXD tables.)
use crate::slb::VoiceLang;
/// The role a cutscene slot plays in the mission flow, from its slot key.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum MovieKind {
/// `LOGO*`, `ADVERTISE_MOVIE`, `STAFF_ROLL` — boot/system movies.
System,
/// `MS<NN><X>` — story intro movie (camera on the speaker; stereo `VOICE_S*`).
Intro,
/// `STAGE<NN>_PHASE0<N>` — in-mission phase cutscene (radio box; `VOICE_RT*`).
Phase,
/// `STAGE<NN>_PHASE_END[_0N]` — phase-end cutscene.
PhaseEnd,
/// `S<NN>_SUPPLY_<ACROPOLIS|TANKER>` — resupply (hokyu) cutscene.
Supply,
}
/// One cutscene slot's manifest row: its semantic slot plus the bound assets.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct MovieEntry {
/// The slot key, e.g. `STAGE01_PHASE01`, `MS00A`, `S02_SUPPLY_ACROPOLIS`.
/// Empty when parsed from a blob without the slot-key array (fallback).
pub slot: String,
/// The slot's role in the mission flow (intro vs. phase cutscene vs. supply …).
pub kind: MovieKind,
/// Mission/stage number (`STAGE<NN>` / `MS<NN>` / `S<NN>_SUPPLY`), if any.
pub mission: Option<u8>,
/// Phase number for [`MovieKind::Phase`]/[`MovieKind::PhaseEnd`], if any.
pub phase: Option<u8>,
/// Movie basename without extension, e.g. `S13A`, `RT01A`, `hokyu_DS_s13A`.
pub movie: String,
/// Voice bank basename (`VOICE_S13A`, `VOICE_D_450`), or `None` for no voice.
pub voice_token: Option<String>,
/// Subtitle table basename (`SUBTITLE_S13A.tbl`), pak prefix stripped, if any.
pub subtitle: Option<String>,
/// On-screen text-overlay (`TELOP`) basename (`pwrt01.prt`), if any.
pub telop: Option<String>,
}
/// Does this blob look like the movie manifest? (`IDXD` whose string pool
/// carries movie + subtitle + voice references.) Used to locate it in
/// `tables.pak` without hardcoding a hash.
pub fn is_manifest(bytes: &[u8]) -> bool {
bytes.len() >= 4
&& &bytes[0..4] == b"IDXD"
&& contains(bytes, b".wmv")
&& contains(bytes, b"VOICE_")
&& contains(bytes, b"SUBTITLE_")
}
/// The field-type keys that appear once in the value region as a schema template.
const FIELD_KEYS: [&str; 4] = ["MOVIE", "VOICETRACK", "SUBTITLE", "TELOP"];
/// Parse the manifest into per-slot cutscene rows, in play order.
///
/// Each returned [`MovieEntry`] carries a movie; valueless slots (e.g. stage 5's
/// absent resupply) are omitted. When the two-array slot/record layout is present
/// (the real manifest) every row is tagged with its authoritative slot/kind/
/// mission/phase; otherwise (a blob without the key array) the kind is inferred
/// from the movie name and `slot` is empty.
pub fn parse(bytes: &[u8]) -> Vec<MovieEntry> {
let toks = ascii_runs(bytes, 3);
// Locate the slot-key array (`LOGO1` … first `<movie>.wmv`) and the value
// region (from the first `.wmv` onward). Fall back to a keyless scrape.
let first_wmv = toks.iter().position(|t| t.ends_with(".wmv"));
let key_start = toks.iter().position(|t| t == "LOGO1");
let (slot_keys, value_toks): (Vec<&str>, &[String]) = match (key_start, first_wmv) {
(Some(k), Some(w)) if k < w => (
toks[k..w].iter().map(String::as_str).collect(),
&toks[w..],
),
_ => (Vec::new(), toks.as_slice()),
};
// Build the value records (drop the schema-template field keys).
let mut records: Vec<MovieEntry> = Vec::new();
for t in value_toks.iter().filter(|t| !FIELD_KEYS.contains(&t.as_str())) {
if let Some(name) = t.strip_suffix(".wmv") {
records.push(MovieEntry {
slot: String::new(),
kind: kind_from_movie(name),
mission: None,
phase: None,
movie: name.to_string(),
voice_token: None,
subtitle: None,
telop: None,
});
} else if let Some(rec) = records.last_mut() {
if let Some(v) = t.strip_prefix("VOICE_") {
if rec.voice_token.is_none() {
rec.voice_token = Some(format!("VOICE_{v}"));
}
} else if t.ends_with(".tbl") {
rec.subtitle.get_or_insert_with(|| after_plus(t));
} else if t.ends_with(".prt") {
rec.telop.get_or_insert_with(|| after_plus(t));
}
}
}
// Attach slot keys to records, recovering valueless-slot gaps via MS anchors.
if !slot_keys.is_empty() {
let ms_targets: std::collections::HashSet<String> = slot_keys
.iter()
.filter_map(|k| ms_target(k))
.collect();
let mut j = 0usize;
for slot in &slot_keys {
if j >= records.len() {
break;
}
// A gap: this slot's "record" is actually the next MS slot's movie.
let would_be = &records[j].movie;
let is_gap = match ms_target(slot) {
// MS slot: only binds if its record is the expected `S<NN><X>`.
Some(expected) => *would_be != expected,
// Non-MS slot: a gap iff the record belongs to an upcoming MS slot.
None => ms_targets.contains(would_be),
};
if is_gap {
continue;
}
let (kind, mission, phase) = classify_slot(slot);
let rec = &mut records[j];
rec.slot = (*slot).to_string();
rec.kind = kind;
rec.mission = mission;
rec.phase = phase;
j += 1;
}
}
records
}
/// If `slot` is an `MS<NN><X>` intro slot, the movie it must bind to (`S<NN><X>`).
fn ms_target(slot: &str) -> Option<String> {
let rest = slot.strip_prefix("MS")?;
// `<NN><X>`: two digits then a letter suffix.
if rest.len() >= 3 && rest[..2].bytes().all(|b| b.is_ascii_digit()) {
Some(format!("S{rest}"))
} else {
None
}
}
/// Semantic role + mission/phase from a slot key.
fn classify_slot(slot: &str) -> (MovieKind, Option<u8>, Option<u8>) {
if let Some(rest) = slot.strip_prefix("MS") {
if rest.len() >= 2 {
return (MovieKind::Intro, rest[..2].parse().ok(), None);
}
}
if let Some(rest) = slot.strip_prefix("STAGE") {
let nn: String = rest.chars().take_while(|c| c.is_ascii_digit()).collect();
if let Ok(m) = nn.parse::<u8>() {
if slot.contains("PHASE_END") {
// trailing `_NN` if present (`STAGE12_PHASE_END_02`)
let phase = slot.rsplit('_').next().and_then(|x| x.parse().ok());
return (MovieKind::PhaseEnd, Some(m), phase);
}
// `STAGE<NN>_PHASE0<N>`
let phase = rest[nn.len()..]
.strip_prefix("_PHASE")
.and_then(|p| p.parse().ok());
return (MovieKind::Phase, Some(m), phase);
}
}
if slot.contains("_SUPPLY") {
if let Some(rest) = slot.strip_prefix('S') {
let nn: String = rest.chars().take_while(|c| c.is_ascii_digit()).collect();
return (MovieKind::Supply, nn.parse().ok(), None);
}
}
(MovieKind::System, None, None)
}
/// Fallback kind from a movie basename (used when no slot-key array is present).
fn kind_from_movie(movie: &str) -> MovieKind {
if movie.starts_with("hokyu_") {
MovieKind::Supply
} else if movie.starts_with("RT") {
MovieKind::Phase
} else if movie.len() >= 3
&& movie.starts_with('S')
&& movie[1..3].bytes().all(|b| b.is_ascii_digit())
{
MovieKind::Intro
} else {
MovieKind::System
}
}
/// The part of a `<pak>+<name>` reference after the `+` (or the whole string).
fn after_plus(s: &str) -> String {
s.rsplit('+').next().unwrap_or(s).to_string()
}
/// The voice bank basename bound to `movie`, if the manifest lists one.
///
/// Returns `None` both when the movie is absent and when it is present with no
/// voice — callers that need to distinguish should use [`parse`].
pub fn voice_token(bytes: &[u8], movie: &str) -> Option<String> {
parse(bytes)
.into_iter()
.find(|e| e.movie == movie)
.and_then(|e| e.voice_token)
}
/// Resolve a movie to its full `sound.pak` voice-entry name for `lang`, using the
/// manifest for the *token* and `sounds.tbl` for the token's *directory* (Movie /
/// etc / Voice differ per token). `None` = the movie has no voice-over.
pub fn resolve_voice_entry(
manifest: &[u8],
sounds_tbl: &[u8],
movie: &str,
lang: VoiceLang,
) -> Option<String> {
let token = voice_token(manifest, movie)?;
resolve_token_path(sounds_tbl, &token, lang)
}
/// Find the full `<lang>\…\<token>.slb` path for a voice `token` in a decompressed
/// `sounds.tbl` (the token's subdir — `Movie`, `etc`, `Voice` — is not fixed).
pub fn resolve_token_path(sounds_tbl: &[u8], token: &str, lang: VoiceLang) -> Option<String> {
let prefix = format!("{}\\", lang.code_pub());
let suffix = format!("\\{token}.slb");
ascii_runs(sounds_tbl, 6)
.into_iter()
.find(|r| r.starts_with(&prefix) && r.ends_with(&suffix))
}
fn contains(hay: &[u8], needle: &[u8]) -> bool {
hay.windows(needle.len()).any(|w| w == needle)
}
/// Collect printable-ASCII runs of at least `min` chars.
fn ascii_runs(bytes: &[u8], min: usize) -> Vec<String> {
let mut out = Vec::new();
let mut cur = String::new();
for &b in bytes {
if (0x20..=0x7e).contains(&b) {
cur.push(b as char);
} else {
if cur.len() >= min {
out.push(std::mem::take(&mut cur));
} else {
cur.clear();
}
}
}
if cur.len() >= min {
out.push(cur);
}
out
}
#[cfg(test)]
mod tests {
use super::*;
fn synth() -> Vec<u8> {
// Minimal IDXD string pool: three movies — a standard one, a hokyu bound
// to a VOICE_D radio clip, and a hokyu with no voice.
let mut b = b"IDXD".to_vec();
b.extend_from_slice(&[0, 0, 0, 0x69, 0, 0]); // binary header (as on disc)
for s in [
"S13A.wmv",
"eng.pak+SUBTITLE_S13A.tbl",
"VOICE_S13A",
"hokyu_LS_s02A.wmv",
"eng.pak+SUBTITLE_hokyu_LS_s02A.tbl",
"VOICE_D_450",
"hokyu_DS_s13A.wmv",
"eng.pak+SUBTITLE_hokyu_DS_s13A.tbl",
] {
b.extend_from_slice(s.as_bytes());
b.push(0);
}
b
}
#[test]
fn groups_movies_and_binds_voice() {
// Keyless fallback path: movie + voice + subtitle + inferred kind.
let m = parse(&synth());
assert_eq!(m.len(), 3);
assert_eq!(m[0].movie, "S13A");
assert_eq!(m[0].voice_token.as_deref(), Some("VOICE_S13A"));
assert_eq!(m[0].kind, MovieKind::Intro);
assert_eq!(m[0].subtitle.as_deref(), Some("SUBTITLE_S13A.tbl"));
assert_eq!(m[1].voice_token.as_deref(), Some("VOICE_D_450"));
assert_eq!(m[1].kind, MovieKind::Supply);
assert_eq!(m[2].voice_token, None, "hokyu_DS_s13A has no voice-over");
assert_eq!(m[2].kind, MovieKind::Supply);
}
/// Build a two-array (keys-then-records) manifest like the real one.
fn synth_two_array(keys: &[&str], records: &[&[&str]]) -> Vec<u8> {
let mut b = b"IDXD".to_vec();
b.extend_from_slice(&[0, 0, 0, 0x05, 0, 0]);
for k in keys {
b.extend_from_slice(k.as_bytes());
b.push(0);
}
for rec in records {
for s in *rec {
b.extend_from_slice(s.as_bytes());
b.push(0);
}
}
b
}
#[test]
fn two_array_slots_kind_mission_phase() {
let blob = synth_two_array(
&["LOGO1", "MS01A", "STAGE01_PHASE01", "STAGE01_PHASE_END_02"],
&[
&["logo1.wmv", "MOVIE"],
&["S01A.wmv", "eng.pak+SUBTITLE_S01A.tbl", "VOICE_S01A"],
&["RT01A.wmv", "eng.pak+pwrt01.prt", "VOICE_RT01A", "VOICETRACK"],
&["RT01C_2.wmv", "VOICE_RT01C_2"],
],
);
let m = parse(&blob);
assert_eq!(m.len(), 4);
assert_eq!((m[0].slot.as_str(), m[0].kind), ("LOGO1", MovieKind::System));
assert_eq!(
(m[1].slot.as_str(), m[1].kind, m[1].mission),
("MS01A", MovieKind::Intro, Some(1))
);
assert_eq!(
(m[2].slot.as_str(), m[2].kind, m[2].mission, m[2].phase),
("STAGE01_PHASE01", MovieKind::Phase, Some(1), Some(1))
);
assert_eq!(m[2].telop.as_deref(), Some("pwrt01.prt"));
assert_eq!(
(m[3].slot.as_str(), m[3].kind, m[3].phase),
("STAGE01_PHASE_END_02", MovieKind::PhaseEnd, Some(2))
);
}
#[test]
fn valueless_supply_slot_is_skipped_via_ms_anchor() {
// S05_SUPPLY has no record; its "record" (S06A) belongs to the next MS slot.
let blob = synth_two_array(
&["LOGO1", "MS05A", "S05_SUPPLY_ACROPOLIS", "MS06A"],
&[
&["logo1.wmv"],
&["S05A.wmv", "VOICE_S05A"],
&["S06A.wmv", "VOICE_S06A"],
],
);
let m = parse(&blob);
assert_eq!(m.len(), 3, "the valueless supply slot yields no entry");
assert_eq!((m[1].slot.as_str(), m[1].movie.as_str()), ("MS05A", "S05A"));
// S06A binds to MS06A, NOT to the intervening valueless supply slot.
assert_eq!((m[2].slot.as_str(), m[2].movie.as_str()), ("MS06A", "S06A"));
assert!(m.iter().all(|e| e.slot != "S05_SUPPLY_ACROPOLIS"));
}
#[test]
fn resolves_token_directory_from_sounds_tbl() {
// sounds.tbl places the two tokens in different subdirs.
let mut tbl = Vec::new();
for s in ["eng\\Movie\\VOICE_S13A.slb", "eng\\etc\\VOICE_D_450.slb"] {
tbl.extend_from_slice(s.as_bytes());
tbl.push(0);
}
let man = synth();
assert_eq!(
resolve_voice_entry(&man, &tbl, "S13A", VoiceLang::English).as_deref(),
Some("eng\\Movie\\VOICE_S13A.slb")
);
assert_eq!(
resolve_voice_entry(&man, &tbl, "hokyu_LS_s02A", VoiceLang::English).as_deref(),
Some("eng\\etc\\VOICE_D_450.slb")
);
assert_eq!(
resolve_voice_entry(&man, &tbl, "hokyu_DS_s13A", VoiceLang::English),
None
);
}
}

View File

@@ -0,0 +1,454 @@
//! Movie cutscene subtitles: the movie → track → text chain.
//!
//! Reverse-engineered statically (see `docs/re/structures/movie-subtitles.md`).
//! A cutscene's on-screen captions are assembled from three places on the disc:
//!
//! 1. `dat/movie/<lang>.pak` holds one **timing track** per movie, keyed by
//! [`track_key`] = `name_hash("subtitle_<movie>.tbl")`. Each track is a
//! `Z1`+zlib block that decompresses to an **IXUD** container whose UTF-16**LE**
//! payload is `SUBTITLE MSG_DEMO_<demo> <mm:ss.cc> …` — i.e. *which*
//! demo-message shows *when* (radio / resupply movies), or inline text.
//! 2. `dat/GP_MAIN_GAME_<L>.pak` holds the **caption text**: more `Z1`+zlib IXUD
//! blocks where each line is stored as `text` immediately followed by its key
//! `MSG_DEMO_<demo>_<page>_<line>` (multi-line captions split across lines).
//!
//! [`load`] joins the two into timed [`SubCue`]s.
//!
//! Note: the IXUD payload here is UTF-16 **little-endian** (verified on the real
//! disc); this is deliberately separate from the older big-endian [`crate::ixud`]
//! presenter, which targets a different (uncompressed) variant.
use std::collections::BTreeMap;
use crate::hash::name_hash;
use crate::pak::PakArchive;
/// Subtitle language. `pak_code` selects `dat/movie/<code>.pak`; `game_code`
/// selects `dat/GP_MAIN_GAME_<code>.pak` (the caption text).
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum SubLang {
English,
Japanese,
German,
French,
Spanish,
Italian,
}
impl SubLang {
/// All languages, in menu order.
pub const ALL: [SubLang; 6] = [
SubLang::English,
SubLang::Japanese,
SubLang::German,
SubLang::French,
SubLang::Spanish,
SubLang::Italian,
];
/// `dat/movie/<code>.pak` stem (the timing tracks + caption font).
pub fn pak_code(self) -> &'static str {
match self {
SubLang::English => "eng",
SubLang::Japanese => "jpn",
SubLang::German => "deu",
SubLang::French => "fra",
SubLang::Spanish => "esp",
SubLang::Italian => "ita",
}
}
/// `dat/GP_MAIN_GAME_<code>.pak` suffix (the caption text pack).
pub fn game_code(self) -> &'static str {
match self {
SubLang::English => "E",
SubLang::Japanese => "J",
SubLang::German => "D",
SubLang::French => "F",
SubLang::Spanish => "S",
SubLang::Italian => "I",
}
}
/// Human label for a UI selector. Kept to Latin script so it renders in a
/// default (non-CJK) UI font; Japanese is romanized for the same reason.
pub fn label(self) -> &'static str {
match self {
SubLang::English => "English",
SubLang::Japanese => "Japanese",
SubLang::German => "Deutsch",
SubLang::French => "Français",
SubLang::Spanish => "Español",
SubLang::Italian => "Italiano",
}
}
}
/// One timed caption line.
#[derive(Debug, Clone, PartialEq)]
pub struct SubCue {
pub start: f32,
pub end: Option<f32>,
pub text: String,
}
/// The `<lang>.pak` TOC key for a movie's subtitle timing track.
pub fn track_key(movie_basename: &str) -> u32 {
name_hash(&format!("subtitle_{movie_basename}.tbl"))
}
/// Load and resolve a movie's timed captions in `lang`.
///
/// `lang_pak` is `dat/movie/<lang>.pak` (+segments); `text_pak` is
/// `dat/GP_MAIN_GAME_<L>.pak` (+segments). Returns the cues in start order, or an
/// empty vec if the movie has no timed track (e.g. story movies whose text is a
/// pre-rendered title card).
pub fn load(movie_basename: &str, lang_pak: &PakArchive, text_pak: &PakArchive) -> Vec<SubCue> {
let Some(Ok(track)) = lang_pak.read_by_hash(track_key(movie_basename)) else {
return Vec::new();
};
if !is_ixud(&track) {
return Vec::new();
}
let raw = parse_track(&track);
if raw.is_empty() {
return Vec::new();
}
// Only build the (large) text table if the track actually references demos.
let needs_text = raw.iter().any(|(t, _, _)| demo_ref(t).is_some());
let text = if needs_text {
build_demo_text(text_pak)
} else {
BTreeMap::new()
};
let mut cues = Vec::new();
for (token, start, end) in raw {
let resolved = match demo_ref(&token) {
Some(demo) => match text.get(&demo) {
Some(lines) => lines.join("\n"),
None => continue, // demo id with no text on disc → skip
},
None => clean(&token), // already inline text
};
if resolved.trim().is_empty() {
continue;
}
cues.push(SubCue {
start,
end,
text: resolved,
});
}
cues.sort_by(|a, b| a.start.partial_cmp(&b.start).unwrap_or(std::cmp::Ordering::Equal));
cues
}
/// The `MSG_DEMO_<id>` ids a movie's timing track references, in track order
/// (empty for inline-text or absent tracks). Language-independent — the same
/// demo ids appear in every `<lang>.pak`. Used to share a voice clip between
/// movies that play the same demo line (the resupply cutscenes reuse one clip
/// per line while only the video differs).
pub fn track_demo_ids(lang_pak: &PakArchive, movie_basename: &str) -> Vec<u32> {
let Some(Ok(track)) = lang_pak.read_by_hash(track_key(movie_basename)) else {
return Vec::new();
};
if !is_ixud(&track) {
return Vec::new();
}
parse_track(&track)
.iter()
.filter_map(|(t, _, _)| demo_ref(t))
.collect()
}
/// Per-line `(demo id, start seconds)` for a movie's timing track — the radio-box
/// voice schedule. Each `MSG_DEMO_<id>` reference takes the start time of the
/// timing that follows it (a caption may list two demo lines under one timing;
/// both then share that start). Empty for inline-text / absent tracks (those are
/// story movies whose voice is the single manifest `VOICETRACK`, not per-line).
pub fn track_voice_cues(lang_pak: &PakArchive, movie_basename: &str) -> Vec<(u32, f32)> {
let Some(Ok(track)) = lang_pak.read_by_hash(track_key(movie_basename)) else {
return Vec::new();
};
if !is_ixud(&track) {
return Vec::new();
}
let toks = utf16le_tokens(&track);
let start = toks
.iter()
.position(|t| t.ends_with("SUBTITLE"))
.map(|i| i + 1)
.unwrap_or(0);
let mut out = Vec::new();
let mut pending: Vec<u32> = Vec::new();
for tok in &toks[start..] {
match parse_timing(tok) {
Some((s, _)) => {
for d in pending.drain(..) {
out.push((d, s));
}
}
None => {
if let Some(d) = demo_ref(tok) {
pending.push(d);
}
}
}
}
out
}
/// Build `demo id → ordered caption lines` from a `GP_MAIN_GAME_<L>` pack.
///
/// Scans every IXUD block, pairing each text token with the immediately
/// following `MSG_DEMO_<demo>_<page>_<line>` key, and groups the lines per demo
/// ordered by (page, line).
pub fn build_demo_text(text_pak: &PakArchive) -> BTreeMap<u32, Vec<String>> {
// demo -> (page,line) -> text
let mut by_demo: BTreeMap<u32, BTreeMap<(u32, u32), String>> = BTreeMap::new();
for entry in text_pak.entries() {
let Ok(bytes) = text_pak.read(entry) else {
continue;
};
if !is_ixud(&bytes) {
continue;
}
let toks = utf16le_tokens(&bytes);
for w in toks.windows(2) {
if let Some((demo, page, line)) = text_key(&w[1]) {
// Pair only when the preceding token is real text — not another
// MSG_DEMO key (bare keys are also serialized consecutively in the
// record directory, which would otherwise masquerade as text).
if demo_ref(&w[0]).is_none()
&& text_key(&w[0]).is_none()
&& !w[0].trim().is_empty()
{
by_demo
.entry(demo)
.or_default()
.insert((page, line), clean(&w[0]));
}
}
}
}
by_demo
.into_iter()
.map(|(d, m)| (d, m.into_values().collect()))
.collect()
}
/// Parse a timing track's IXUD payload into `(token, start, end)` triples, where
/// `token` is either a `MSG_DEMO_<d>` reference or an inline caption string.
///
/// A single on-screen caption may be stored as **several consecutive text
/// tokens** followed by one timing (the source hard-splits a multi-line caption
/// into one token per line — e.g. `"Look at it father"` + `"& beautiful isn't
/// it"` share the `01:14.80-01:17.60` timing). We therefore accumulate every
/// text token seen since the last timing and join them with a newline when the
/// timing arrives, instead of pairing strictly 1:1 (which silently dropped every
/// line but the last of a multi-line caption).
fn parse_track(ixud: &[u8]) -> Vec<(String, f32, Option<f32>)> {
let toks = utf16le_tokens(ixud);
let start = toks
.iter()
.position(|t| t.ends_with("SUBTITLE"))
.map(|i| i + 1)
.unwrap_or(0);
let rest = &toks[start..];
let mut out = Vec::new();
let mut pending: Vec<&str> = Vec::new();
for tok in rest {
match parse_timing(tok) {
Some((s, e)) => {
if !pending.is_empty() {
out.push((pending.join("\n"), s, e));
pending.clear();
}
}
None => pending.push(tok),
}
}
out
}
fn is_ixud(b: &[u8]) -> bool {
b.len() >= 4 && &b[0..4] == b"IXUD"
}
/// Normalize a caption string: the source stores line breaks as the literal
/// two-character escape `\n`; turn it into a real newline and trim.
fn clean(s: &str) -> String {
s.replace("\\n", "\n").trim().to_string()
}
/// `MSG_DEMO_<d>` → `Some(d)` (reference token, no page/line).
fn demo_ref(t: &str) -> Option<u32> {
let rest = t.strip_prefix("MSG_DEMO_")?;
if rest.contains('_') {
return None; // that's a text key (has page/line), not a bare reference
}
rest.parse().ok()
}
/// `MSG_DEMO_<d>_<page>_<line>` → `Some((d,page,line))` (text key).
fn text_key(t: &str) -> Option<(u32, u32, u32)> {
let rest = t.strip_prefix("MSG_DEMO_")?;
let mut it = rest.split('_');
let d = it.next()?.parse().ok()?;
let p = it.next()?.parse().ok()?;
let l = it.next()?.parse().ok()?;
if it.next().is_some() {
return None;
}
Some((d, p, l))
}
/// Extract UTF-16**LE** text runs from a blob, **independent of byte alignment**.
///
/// The IXUD string pool packs records at odd byte offsets, so a fixed 2-byte
/// stride from the block start misreads every character. Instead we scan a
/// sliding window, decoding each 16-bit LE unit and keeping runs of "text-like"
/// code points (ASCII, Latin-1 supplement — the German umlauts / accented
/// Latin — and Latin Extended-A). A non-text unit ends the run and we advance by
/// one byte to re-lock onto the correct parity of the next run. Accepting the
/// Latin ranges (not just ASCII) is what keeps `müsste`, `Français`, `español`
/// intact — earlier the accented char *and its predecessor* were dropped.
///
/// CJK (Japanese) is deliberately out of range: those units also occur as noise
/// at the wrong parity, and the UI font can't render them anyway.
fn utf16le_tokens(bytes: &[u8]) -> Vec<String> {
let mut tokens = Vec::new();
let mut cur = String::new();
let mut i = 0;
while i + 1 < bytes.len() {
let u = u16::from_le_bytes([bytes[i], bytes[i + 1]]);
if is_text_unit(u) {
if let Some(c) = char::from_u32(u as u32) {
cur.push(c);
}
i += 2;
} else {
if cur.chars().count() >= 2 {
tokens.push(std::mem::take(&mut cur));
} else {
cur.clear();
}
i += 1;
}
}
if cur.chars().count() >= 2 {
tokens.push(cur);
}
tokens
}
/// Whether a UTF-16 unit is a printable Latin text character we keep in a run.
/// Excludes the C0/C1 control blocks (incl. 0x7F0x9F) so binary noise breaks
/// runs instead of joining them.
fn is_text_unit(u: u16) -> bool {
matches!(u, 0x20..=0x7e | 0xa0..=0xff | 0x100..=0x17f)
}
/// Parse `MM:SS.ss` or a `start-end` range into seconds.
fn parse_timing(s: &str) -> Option<(f32, Option<f32>)> {
let one = |p: &str| -> Option<f32> {
let (mm, ss) = p.trim().split_once(':')?;
let m: f32 = if mm.trim().is_empty() {
0.0
} else {
mm.trim().parse().ok()?
};
let sec: f32 = ss.trim().parse().ok()?;
Some(m * 60.0 + sec)
};
match s.split_once('-') {
Some((a, b)) => Some((one(a)?, Some(one(b)?))),
None => Some((one(s)?, None)),
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn track_key_matches_disc() {
// Verified against dat/movie/eng.pak.
assert_eq!(track_key("S00A"), 0x6F2D_9663);
assert_eq!(track_key("hokyu_DS_s02A"), 0x3662_B1F8);
assert_eq!(track_key("RT01C_1"), 0x756F_69FB);
}
#[test]
fn demo_and_text_keys() {
assert_eq!(demo_ref("MSG_DEMO_192"), Some(192));
assert_eq!(demo_ref("MSG_DEMO_604_000_00"), None);
assert_eq!(text_key("MSG_DEMO_604_000_01"), Some((604, 0, 1)));
assert_eq!(text_key("MSG_DEMO_192"), None);
}
#[test]
fn keeps_latin1_accents() {
// "müsste" must survive intact (umlaut + its neighbour), not collapse to
// "sste". Encode as UTF-16LE at an ODD start offset to exercise the
// alignment-independent scan.
let mut b = vec![0xAAu8]; // 1 junk byte → strings start at an odd offset
for s in ["müsste", "Français", "español"] {
for u in s.encode_utf16() {
b.extend_from_slice(&u.to_le_bytes());
}
b.extend_from_slice(&[0, 0]);
}
let toks = utf16le_tokens(&b);
assert!(toks.contains(&"müsste".to_string()), "got {toks:?}");
assert!(toks.contains(&"Français".to_string()), "got {toks:?}");
assert!(toks.contains(&"español".to_string()), "got {toks:?}");
}
/// Build a synthetic UTF-16LE IXUD and check the LE token walk + timing pair.
#[test]
fn parses_le_track() {
let mut b = b"IXUD".to_vec();
b.extend_from_slice(&[0, 1, 0, 0, 0x70, 0x3E, 0xC8, 0x6C]);
for t in ["SUBTITLE", "MSG_DEMO_5", "00:01.50", "Hello", "00:03.00"] {
for u in t.encode_utf16() {
b.extend_from_slice(&u.to_le_bytes());
}
b.extend_from_slice(&[0, 0]);
}
let raw = parse_track(&b);
assert_eq!(raw.len(), 2);
assert_eq!(raw[0], ("MSG_DEMO_5".to_string(), 1.5, None));
assert_eq!(raw[1], ("Hello".to_string(), 3.0, None));
}
/// Two text tokens before a single timing = one two-line caption; the first
/// line must NOT be dropped (the real S13A `Look at it father` regression).
#[test]
fn joins_multiline_caption() {
let mut b = b"IXUD".to_vec();
b.extend_from_slice(&[0, 1, 0, 0, 0x70, 0x3E, 0xC8, 0x6C]);
for t in [
"SUBTITLE",
"That is such magnificent power.",
"01:09.70-01:12.20",
"Look at it father",
"& beautiful isn't it",
"01:14.80-01:17.60",
] {
for u in t.encode_utf16() {
b.extend_from_slice(&u.to_le_bytes());
}
b.extend_from_slice(&[0, 0]);
}
let raw = parse_track(&b);
assert_eq!(raw.len(), 2);
assert_eq!(raw[0].0, "That is such magnificent power.");
assert_eq!(
raw[1].0, "Look at it father\n& beautiful isn't it",
"both lines of the caption must be kept"
);
assert_eq!((raw[1].1, raw[1].2), (74.8, Some(77.6)));
}
}

View File

@@ -0,0 +1,142 @@
//! Cutscene-voice resolution: movie → sound-id → continuous-stream byte region.
//!
//! Reverse-engineered 2026-07-22 (Canary file-I/O trace, user-confirmed by ear).
//! The movie voices form **one continuous XMA stream**, physically chunked into
//! `<lang>\Movie\VOICE_*.slb` `sound.pak` TOC entries whose boundaries do **not**
//! align with the cutscene cues — a single cue routinely spans two `.slb` chunks,
//! so a `.slb` does not necessarily hold the track its name claims. Each cue's
//! audio ends at an inline **trailer descriptor** `(sound_id: u32be, 0x11, …)`
//! whose id repeats at +0x800. Cue N's audio = the stream bytes
//! `[descriptor(N-1) .. descriptor(N)]`, read continuously across chunk
//! boundaries. Verified: decoded voice length matches each movie within ~0.3 s,
//! including cross-boundary cues (RT01B/RT01C span `VOICE_ADV`→`VOICE_RT01A`→…).
//!
//! Resolution chain:
//! 1. movie → cue name (`ADVERTISE_MOVIE` manifest `VOICETRACK`, e.g. `VOICE_RT01A`)
//! 2. cue name → sound-id (master sound registry, [`registry_voice_ids`], → 1601)
//! 3. sound-id → `[start, end]` global offsets via [`find_descriptor`] over the
//! `sound.pNN` stream (the caller supplies the bytes + the physical anchor).
use std::collections::HashMap;
/// Parse the master sound registry — the large per-language `tables.pak` IDXD
/// entry (schema `0x13cb84ba`) that carries the `<lang>\Movie\VOICE_*.slb` paths —
/// into `cue-name → sound-id` from its `NAME\0 ID\0` string pool
/// (e.g. `VOICE_ADV`→1600, `VOICE_RT01A`→1601).
pub fn registry_voice_ids(registry: &[u8]) -> HashMap<String, u32> {
let mut toks: Vec<&[u8]> = Vec::new();
let mut start = 0usize;
for i in 0..registry.len() {
if registry[i] == 0 {
if i > start {
toks.push(&registry[start..i]);
}
start = i + 1;
}
}
let mut out = HashMap::new();
for w in toks.windows(2) {
let (name, num) = (w[0], w[1]);
if name.starts_with(b"VOICE_")
&& !num.is_empty()
&& num.iter().all(u8::is_ascii_digit)
{
if let (Ok(n), Ok(id)) = (
std::str::from_utf8(name),
std::str::from_utf8(num).unwrap().parse::<u32>(),
) {
out.entry(n.to_string()).or_insert(id);
}
}
}
out
}
/// The constant marker following a cue's sound-id in its inline trailer.
const DESC_MARK: u32 = 0x11;
/// The trailer repeats its id this many bytes later; used to reject a coincidental
/// `(id, 0x11)` pair inside XMA audio (probability of a false match is ~2⁻⁶⁴).
const DESC_REPEAT: usize = 0x800;
/// Byte offset (within `buf`, a slice of the continuous voice stream) of cue
/// `id`'s trailer descriptor: the first big-endian `(id, 0x11)` pair whose id
/// repeats at `+0x800`. `None` if absent in the slice.
pub fn find_descriptor(buf: &[u8], id: u32) -> Option<usize> {
if buf.len() < DESC_REPEAT + 4 {
return None;
}
let be = |o: usize| u32::from_be_bytes([buf[o], buf[o + 1], buf[o + 2], buf[o + 3]]);
let end = buf.len() - (DESC_REPEAT + 4);
let mut o = 0;
while o <= end {
if be(o) == id && be(o + 4) == DESC_MARK && be(o + DESC_REPEAT) == id {
return Some(o);
}
o += 4;
}
None
}
/// Plausible sound-id range for a trailer (all real cue ids are well under this).
const ID_MAX: u32 = 0x1_0000;
/// Offset of the nearest cue trailer strictly **before** `before` — any plausible
/// id. Used to bound a cue's START when its `id-1` predecessor doesn't exist
/// (the sound-id sequence has gaps, e.g. `VOICE_D_453`=7142 → `VOICE_D_454`=7188).
/// Scans downward and returns the first (largest-offset) valid `(id, 0x11)` with
/// the `+0x800` id-repeat. `None` if there is no trailer below `before`.
pub fn find_descriptor_before(buf: &[u8], before: usize) -> Option<usize> {
if before < 4 {
return None;
}
let cap = buf.len().checked_sub(DESC_REPEAT + 4)?;
let be = |o: usize| u32::from_be_bytes([buf[o], buf[o + 1], buf[o + 2], buf[o + 3]]);
// Start strictly below `before` (4-aligned) so the target's own trailer is
// never returned as its predecessor.
let mut o = (before - 1).min(cap) & !3;
loop {
let id = be(o);
if id >= 1 && id < ID_MAX && be(o + 4) == DESC_MARK && be(o + DESC_REPEAT) == id {
return Some(o);
}
if o < 4 {
return None;
}
o -= 4;
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn registry_parses_name_id_pairs() {
let mut b = Vec::new();
for tok in [
b"VOICE_ADV".as_slice(),
b"1600",
b"VOICE_RT01A",
b"1601",
b"NOT_A_VOICE",
b"9",
] {
b.extend_from_slice(tok);
b.push(0);
}
let ids = registry_voice_ids(&b);
assert_eq!(ids.get("VOICE_ADV"), Some(&1600));
assert_eq!(ids.get("VOICE_RT01A"), Some(&1601));
assert_eq!(ids.get("NOT_A_VOICE"), None);
}
#[test]
fn find_descriptor_matches_repeat() {
let mut b = vec![0u8; 0x900];
b[0..4].copy_from_slice(&1601u32.to_be_bytes());
b[4..8].copy_from_slice(&0x11u32.to_be_bytes());
b[DESC_REPEAT..DESC_REPEAT + 4].copy_from_slice(&1601u32.to_be_bytes());
assert_eq!(find_descriptor(&b, 1601), Some(0));
assert_eq!(find_descriptor(&b, 1600), None);
}
}

View File

@@ -178,6 +178,18 @@ impl PakArchive {
})
}
/// Parse **only** the TOC from a `.pak` index, without any segment data.
///
/// For huge archives (`sound.pak` is ~1 GB across `.p00..p04`) this lets a
/// caller locate one entry's `(offset, size)` and read just that byte range
/// from the segments, instead of loading the whole thing into memory. The
/// returned entries are sorted by `name_hash` (binary-searchable).
pub fn parse_toc(index: &[u8]) -> Result<Vec<PakEntry>, PakError> {
// Reuse from_parts' validation with an empty data blob; the entries are
// independent of the data.
Ok(Self::from_parts(index, Vec::new())?.entries)
}
/// All TOC entries, in stored order (ascending `name_hash`).
pub fn entries(&self) -> &[PakEntry] {
&self.entries

View File

@@ -0,0 +1,370 @@
//! `.slb` XACT sound banks → a decodable XMA1 `RIFF`.
//!
//! `dat/sound.pak` (9519 entries across `sound.p00..p04`) is the game's audio
//! bank. Each entry is an XACT `.slb` wrapping **XMA1** (`fmt ` tag `0x0165`,
//! 48 kHz). Files are named `<lang>\Voice\…`, `<lang>\Movie\VOICE_<movie>.slb`,
//! `BGM_###.slb`, etc. (see `docs/re/structures/sound-slb.md`); look them up by
//! [`crate::hash::name_hash`]. Movie cutscene voice = one continuous
//! `<eng|jpn>\Movie\VOICE_<movie>.slb` track meant to play from the video start.
//!
//! Two on-disc layouts, both reversed statically:
//! - **RIFF present** — a standard `RIFF/WAVE` sits inside the bank; its 32-byte
//! `fmt ` is the real `XMAWAVEFORMAT` and the XMA packets are everything after
//! that RIFF's `data` chunk header (the declared `data` size is unreliable, so
//! we take to end and let the caller clamp to the known media length).
//! - **Headerless** — no RIFF at all; a fixed **1392-byte** header precedes raw
//! XMA1 packets (48 kHz, 2 channels).
//!
//! [`to_xma_riff`] rebuilds a standalone `RIFF/WAVE` (XMA1) for either layout,
//! ready to hand to an XMA decoder (e.g. FFmpeg's `xma1`). The content is always
//! mono (some clips put it in the left channel only, others duplicate L=R), so
//! the decode step should downmix to mono (take the left channel).
/// Fixed offset of the raw XMA1 stream in a headerless `.slb` (no `RIFF`).
pub const HEADERLESS_DATA_OFFSET: usize = 1392;
/// Voice language for cutscene audio. Only English and Japanese voice exist on
/// the disc (subtitles cover more languages, voice does not).
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum VoiceLang {
English,
Japanese,
}
impl VoiceLang {
fn code(self) -> &'static str {
match self {
VoiceLang::English => "eng",
VoiceLang::Japanese => "jpn",
}
}
/// `eng` / `jpn` — the `sound.pak` path prefix for this voice language.
pub fn code_pub(self) -> &'static str {
self.code()
}
}
/// The `sound.pak` entry name for a movie's continuous voice track, e.g.
/// `eng\Movie\VOICE_RT07A.slb`. Hash it with [`crate::hash::name_hash`] to get
/// the `sound.pak` TOC key.
pub fn movie_voice_name(movie_basename: &str, lang: VoiceLang) -> String {
format!("{}\\Movie\\VOICE_{}.slb", lang.code(), movie_basename)
}
/// One playable voice clip discovered in `sounds.tbl`.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct VoiceClip {
/// Full `sound.pak` entry name, e.g. `eng\Voice\VOICE_ADAN_010.slb`.
pub name: String,
/// Speaker/category code, e.g. `ADAN`, `TCAF`, `A`, `RT07A`.
pub speaker: String,
/// Short UI label, e.g. `ADAN 010`.
pub display: String,
}
/// Enumerate the voice/dialog clips named in a decompressed `sounds.tbl` (the
/// IDXD in `tables.pak`). Extracts every `<lang>\{Voice,etc,Movie,Briefing}\…`
/// path ending in `.slb` for `lang`, parsed into `(name, speaker, display)`.
pub fn list_voice_clips(sounds_tbl: &[u8], lang: VoiceLang) -> Vec<VoiceClip> {
let prefix = format!("{}\\", lang.code());
let mut seen = std::collections::BTreeSet::new();
let mut out = Vec::new();
// Scan for printable-ASCII runs; keep those that look like a voice path.
let mut i = 0;
while i < sounds_tbl.len() {
let start = i;
while i < sounds_tbl.len() && (0x20..=0x7e).contains(&sounds_tbl[i]) {
i += 1;
}
if i - start >= 6 {
if let Ok(s) = std::str::from_utf8(&sounds_tbl[start..i]) {
// Every spoken-line category, so the standalone player covers them
// all: in-mission radio (`\Voice\`, `\etc\`) and bound movie voices
// (`\Movie\`) all carry `VOICE_`; mission-briefing lines live in
// `\Briefing\` as `BR<NN>_<MM>.slb` (no `VOICE` in the name).
let is_voice = s.contains("VOICE") || s.contains("\\Briefing\\");
if s.starts_with(&prefix) && s.ends_with(".slb") && is_voice {
if seen.insert(s.to_string()) {
out.push(parse_voice_clip(s));
}
}
}
}
i += 1;
}
out
}
fn parse_voice_clip(name: &str) -> VoiceClip {
// `<lang>\<cat>\VOICE_<SPK>_<NNN>.slb` or `..\VOICE_<movie>.slb`.
let stem = name
.rsplit('\\')
.next()
.unwrap_or(name)
.strip_suffix(".slb")
.unwrap_or(name);
let body = stem.strip_prefix("VOICE_").unwrap_or(stem);
let (speaker, display) = match body.rsplit_once('_') {
Some((spk, num)) if num.chars().all(|c| c.is_ascii_digit()) => {
(spk.to_string(), format!("{spk} {num}"))
}
_ => (body.to_string(), body.to_string()),
};
VoiceClip {
name: name.to_string(),
speaker,
display,
}
}
/// Build one standalone XMA1 `RIFF/WAVE` **per sub-wave** in a `.slb` bank, in
/// on-disc order. A `.slb` is an XACT bank of one or more sub-waves; the sub-waves
/// are EITHER alternate takes (the first is the whole track — e.g. `VOICE_S01A`,
/// `VOICE_ADV`: sub0 ≈ the movie length) OR sequential segments that must be
/// concatenated (e.g. `VOICE_RT07A`: 24s + 14s + 11s ≈ the 50s movie). The caller
/// decodes each to PCM, concatenates in order, and clamps to the movie length —
/// that yields the full track for segment banks while the clamp drops the
/// duplicate takes for alternate-take banks. (Dynamic RE via Canary file-I/O
/// tracing confirmed the movie→voice binding; this fixes the *decode* of `RT*`.)
pub fn to_xma_riffs(slb: &[u8]) -> Vec<Vec<u8>> {
let mut out = Vec::new();
if find(slb, b"RIFF", 0).is_none() {
// Headerless single-stream bank.
if let Some(data) = slb.get(HEADERLESS_DATA_OFFSET..) {
if !data.is_empty() {
out.push(build_riff(&synth_xma1_fmt(2, 2, 48000), data));
}
}
return out;
}
let mut pos = 0usize;
while let Some(ri) = find(slb, b"RIFF", pos) {
// Parse this sub-wave's fmt + data (declared size is honest per sub-wave).
let Some(fi) = find(slb, b"fmt ", ri) else { break };
let Some(fsz) = le32(slb, fi + 4) else { break };
let Some(fmt_end) = fi.checked_add(8).and_then(|v| v.checked_add(fsz as usize)) else {
break;
};
if fmt_end > slb.len() {
break;
}
let Some(di) = find(slb, b"data", fi) else { break };
let Some(dsz) = le32(slb, di + 4) else { break };
let Some(ds) = di.checked_add(8) else { break };
let de = ds
.checked_add(dsz as usize)
.unwrap_or(slb.len())
.min(slb.len());
if let Some(data) = slb.get(ds..de) {
if !data.is_empty() {
out.push(build_riff(&slb[fi..fmt_end], data));
}
}
// Advance past this sub-wave's data to find the next RIFF.
pos = de.max(ri + 4);
}
out
}
/// Build a standalone XMA1 `RIFF/WAVE` from a `.slb` bank, decodable by FFmpeg's
/// `xma1`. Returns `None` if the bank is too small / malformed. This is the
/// FIRST sub-wave only; prefer [`to_xma_riffs`] for correct multi-segment banks.
pub fn to_xma_riff(slb: &[u8]) -> Option<Vec<u8>> {
if let Some(ri) = find(slb, b"RIFF", 0) {
// RIFF layout: a `.slb` is an XACT bank of one or more sub-waves, each
// `[seek][RIFF: fmt + Dmmy pad + data][declared_size XMA bytes]`. Take the
// FIRST sub-wave, bounded by its **declared `data` size** (which is
// honest per sub-wave). Decoding to end-of-file instead would append the
// later sub-waves — for multi-take story movies those are ALTERNATE takes,
// which is what made S10S16 play the wrong audio.
let fi = find(slb, b"fmt ", ri)?;
let fsz = le32(slb, fi + 4)? as usize;
let fmt_end = fi.checked_add(8)?.checked_add(fsz)?;
if fmt_end > slb.len() {
return None;
}
let fmt_chunk = &slb[fi..fmt_end];
let di = find(slb, b"data", fi)?;
let dsz = le32(slb, di + 4)? as usize;
let end = di.checked_add(8)?.checked_add(dsz)?.min(slb.len());
let data = slb.get(di + 8..end)?;
Some(build_riff(fmt_chunk, data))
} else {
// Headerless: 1392-byte header, then raw XMA1 (48 kHz, 2 channels).
let data = slb.get(HEADERLESS_DATA_OFFSET..)?;
if data.is_empty() {
return None;
}
Some(build_riff(&synth_xma1_fmt(2, 2, 48000), data))
}
}
/// Rebuild a standalone XMA1 `RIFF/WAVE` from a single-stream `.slb`, robust to
/// the layout variants seen in `<lang>\etc\` radio clips. Unlike [`to_xma_riff`]
/// (which assumes `RIFF → fmt → data` in order), this picks the **largest `data`
/// chunk anywhere** in the bank — some radio banks store the audio *before* the
/// trailing `RIFF`/`fmt` metadata (an empty post-`RIFF` `data` chunk), which the
/// ordered scan misses. Pairs it with the first `fmt ` chunk; falls back to the
/// headerless layout. Returns `None` only when no usable audio can be found.
pub fn to_xma_riff_best(slb: &[u8]) -> Option<Vec<u8>> {
// Largest usable `data` chunk (bounded by its declared size and the buffer).
let mut best: Option<(usize, usize)> = None; // (data offset, usable payload len)
let mut i = 0;
while let Some(di) = find(slb, b"data", i) {
let declared = le32(slb, di + 4).unwrap_or(0) as usize;
let usable = declared.min(slb.len().saturating_sub(di + 8));
if best.map_or(true, |(_, b)| usable > b) {
best = Some((di, usable));
}
i = di + 4;
}
if let (Some(fi), Some((di, sz))) = (find(slb, b"fmt ", 0), best) {
if sz > 512 {
let fsz = le32(slb, fi + 4)? as usize;
let fmt_end = (fi + 8 + fsz).min(slb.len());
let data = slb.get(di + 8..di + 8 + sz)?;
return Some(build_riff(slb.get(fi..fmt_end)?, data));
}
}
// Headerless fallback: fixed data offset, synthesized XMA1 stereo/48k fmt.
let data = slb.get(HEADERLESS_DATA_OFFSET..)?;
(!data.is_empty()).then(|| build_riff(&synth_xma1_fmt(2, 2, 48000), data))
}
/// A minimal `fmt ` chunk carrying an XMA1 `XMAWAVEFORMAT` (one stream).
fn synth_xma1_fmt(channels: u8, channel_mask: u16, rate: u32) -> Vec<u8> {
let mut fmt = Vec::with_capacity(40);
fmt.extend_from_slice(b"fmt ");
fmt.extend_from_slice(&32u32.to_le_bytes());
// XMAWAVEFORMAT header
fmt.extend_from_slice(&0x0165u16.to_le_bytes()); // wFormatTag = XMA1
fmt.extend_from_slice(&16u16.to_le_bytes()); // BitsPerSample
fmt.extend_from_slice(&0u16.to_le_bytes()); // EncodeOptions
fmt.extend_from_slice(&0u16.to_le_bytes()); // LargestSkip
fmt.extend_from_slice(&1u16.to_le_bytes()); // NumStreams
fmt.push(0); // LoopCount
fmt.push(3); // Version
// XMASTREAMFORMAT[0]
fmt.extend_from_slice(&(rate * channels as u32 * 2).to_le_bytes()); // PsuedoBytesPerSec
fmt.extend_from_slice(&rate.to_le_bytes()); // SampleRate
fmt.extend_from_slice(&0u32.to_le_bytes()); // LoopStart
fmt.extend_from_slice(&0u32.to_le_bytes()); // LoopEnd
fmt.push(4); // SubframeData
fmt.push(channels); // Channels
fmt.extend_from_slice(&channel_mask.to_le_bytes()); // ChannelMask
fmt
}
fn build_riff(fmt_chunk: &[u8], data: &[u8]) -> Vec<u8> {
let mut body = Vec::with_capacity(4 + fmt_chunk.len() + 8 + data.len());
body.extend_from_slice(b"WAVE");
body.extend_from_slice(fmt_chunk);
body.extend_from_slice(b"data");
body.extend_from_slice(&(data.len() as u32).to_le_bytes());
body.extend_from_slice(data);
let mut out = Vec::with_capacity(8 + body.len());
out.extend_from_slice(b"RIFF");
out.extend_from_slice(&(body.len() as u32).to_le_bytes());
out.extend_from_slice(&body);
out
}
fn find(hay: &[u8], needle: &[u8], from: usize) -> Option<usize> {
if from >= hay.len() {
return None;
}
hay[from..]
.windows(needle.len())
.position(|w| w == needle)
.map(|p| p + from)
}
fn le32(b: &[u8], o: usize) -> Option<u32> {
let s = b.get(o..o + 4)?;
Some(u32::from_le_bytes([s[0], s[1], s[2], s[3]]))
}
#[cfg(test)]
mod tests {
use super::*;
use crate::hash::name_hash;
#[test]
fn movie_voice_names_and_hashes() {
assert_eq!(
movie_voice_name("RT07A", VoiceLang::English),
"eng\\Movie\\VOICE_RT07A.slb"
);
// Verified present in dat/sound.pak (VOICE_S00A == TOC entry 0x2A4F97D3).
assert_eq!(name_hash("eng\\Movie\\VOICE_RT07A.slb"), 0xA44A_EA1C);
assert_eq!(name_hash("eng\\Movie\\VOICE_S00A.slb"), 0x2A4F_97D3);
}
#[test]
fn takes_only_first_subwave() {
// Bank of TWO sub-waves; only the FIRST must be extracted (bounded by its
// declared `data` size), not everything to end-of-file.
let fmt = synth_xma1_fmt(2, 2, 48000);
let mut slb = vec![0xAB; 16];
slb.extend_from_slice(b"RIFF");
slb.extend_from_slice(&999u32.to_le_bytes()); // riff size (ignored)
slb.extend_from_slice(b"WAVE");
slb.extend_from_slice(&fmt);
slb.extend_from_slice(b"data");
slb.extend_from_slice(&4u32.to_le_bytes()); // declared: 4 bytes
slb.extend_from_slice(&[1, 2, 3, 4]); // sub-wave 1 XMA
slb.extend_from_slice(&[9, 9, 9, 9]); // a second sub-wave's bytes — excluded
let riff = to_xma_riff(&slb).unwrap();
assert_eq!(&riff[0..4], b"RIFF");
let dpos = find(&riff, b"data", 0).unwrap();
assert_eq!(le32(&riff, dpos + 4), Some(4));
assert_eq!(&riff[dpos + 8..], &[1, 2, 3, 4]);
}
#[test]
fn list_voice_clips_covers_movie_radio_and_briefing() {
// The standalone player must enumerate every spoken-line category: bound
// movie voices, in-mission radio (\etc\ + \Voice\), and briefing (\Briefing\,
// whose BR<NN>_<MM> names lack "VOICE"). Music (BGM_*) must stay excluded.
let mut tbl = Vec::new();
for s in [
"eng\\Movie\\VOICE_S13A.slb",
"eng\\etc\\VOICE_D_450.slb",
"eng\\Voice\\VOICE_ADAN_010.slb",
"eng\\Briefing\\BR01_01.slb",
"eng\\bgm\\BGM_001.slb",
] {
tbl.extend_from_slice(s.as_bytes());
tbl.push(0);
}
let names: Vec<String> = list_voice_clips(&tbl, VoiceLang::English)
.into_iter()
.map(|c| c.name)
.collect();
for want in [
"eng\\Movie\\VOICE_S13A.slb",
"eng\\etc\\VOICE_D_450.slb",
"eng\\Voice\\VOICE_ADAN_010.slb",
"eng\\Briefing\\BR01_01.slb",
] {
assert!(names.iter().any(|n| n == want), "missing {want}");
}
assert!(
!names.iter().any(|n| n.contains("BGM_")),
"music must not be listed as a voice clip"
);
}
#[test]
fn rebuilds_riff_from_headerless() {
let mut slb = vec![0u8; HEADERLESS_DATA_OFFSET];
slb.extend_from_slice(&[9, 8, 7, 6]);
let riff = to_xma_riff(&slb).unwrap();
let dpos = find(&riff, b"data", 0).unwrap();
assert_eq!(&riff[dpos + 8..], &[9, 8, 7, 6]);
// synthetic fmt advertises XMA1.
let fpos = find(&riff, b"fmt ", 0).unwrap();
assert_eq!(le32(&riff, fpos + 8).map(|v| v as u16), Some(0x0165));
}
}

View File

@@ -1,27 +1,25 @@
//! `T8aD` — the game's 2D UI/HUD texture format.
//!
//! A linear (untiled) 32bpp surface stored **A8R8G8B8** (Xbox byte order), with
//! a small fixed-per-variant header:
//! A 32bpp **A8R8G8B8** (Xbox byte order) surface stored as **256×256 raster
//! tiles in row-major order** — each tile prefixed by a 16-byte tile header, edge
//! tiles clipped to the image bounds. Fully reversed 2026-07-17 from the file
//! header and verified against the running game (title screen).
//!
//! ```text
//! 0x00 4 Magic "T8aD"
//! 0x14 4 width (BE u32)
//! 0x18 4 height (BE u32)
//! 0x1c 4 type field → header size: 1→64 2→84 3→104 4→124 15→344
//! <header> w*h*4 bytes of A8R8G8B8 pixel data, row-major
//! 0x00 4 Magic "T8aD"
//! 0x14 4 width (BE u32)
//! 0x18 4 height (BE u32)
//! 0x1c 4 tile count (BE u32) = ceil(w/256) * ceil(h/256)
//! 0x2c tiles*4 offset table: absolute byte offset of each row-major tile
//! <off> 16 per-tile header (flags + tile w/h), then:
//! <off+16> tile_w * tile_h * 4 bytes of A8R8G8B8 pixels, row-major
//! ```
//!
//! Static forensics over the disc showed ~85% of entries decode exactly with
//! this rule; the rest (unknown type field or a `w*h*4` that doesn't fit — most
//! likely DXT / palettized variants) return `None` rather than a wrong image.
//!
//! Keying the header size off the type field (not `len - w*h*4`) makes the
//! decoder **container-safe**: RATC/LSTA hand us a child slice that may have
//! trailing padding before the next child, so we must not infer the header from
//! the slice length.
//!
//! ⚠️ Colour correctness (channel order / endianness / sRGB) is unverified
//! against the running game — see the RE backlog.
//! Surfaces ≤256px wide are a single tile column, so the first tile's pixels sit
//! at `44 + tiles*4 + 16 = 64` — which is why the old "type→header size 64/84/…"
//! rule (header = 44 + tiles*20) happened to decode small textures correctly: for
//! one tile it lands on the same pixel start. Wide textures were garbled because
//! the offset table + 16-byte per-tile headers weren't accounted for.
/// Magic at the start of every T8aD surface.
pub const T8AD_MAGIC: [u8; 4] = *b"T8aD";
@@ -39,54 +37,74 @@ pub fn is_t8ad(bytes: &[u8]) -> bool {
bytes.len() >= 4 && bytes[0..4] == T8AD_MAGIC
}
/// Header size in bytes for a given type field (`@0x1c`), or `None` for an
/// unrecognized variant.
fn header_size(type_field: u32) -> Option<usize> {
match type_field {
1 => Some(64),
2 => Some(84),
3 => Some(104),
4 => Some(124),
15 => Some(344),
_ => None,
}
}
#[inline]
fn be32(b: &[u8], off: usize) -> u32 {
u32::from_be_bytes([b[off], b[off + 1], b[off + 2], b[off + 3]])
}
/// Side of the square storage tile, in texels, and the per-tile header size.
const TILE: usize = 256;
const TILE_HDR: usize = 16;
/// Decode a T8aD surface from a slice whose first bytes ARE the magic. Returns
/// `None` for non-T8aD input or a variant we can't decode as RGBA (never guesses).
///
/// Layout (reversed from the header + verified against the running game):
/// a 44-byte base header, then a `tiles`-entry big-endian u32 **offset table** at
/// `0x2c`, where `tiles` = the field at `0x1c` = `ceil(w/256) * ceil(h/256)`.
/// Each entry is the absolute byte offset of a **row-major** 256×256 tile; every
/// tile is a 16-byte tile header followed by `tile_w*tile_h*4` A8R8G8B8 pixels,
/// edge tiles clipped to the image bounds.
pub fn parse(bytes: &[u8]) -> Option<T8adImage> {
if !is_t8ad(bytes) || bytes.len() < 0x40 {
return None;
}
let width = be32(bytes, 0x14);
let height = be32(bytes, 0x18);
let width = be32(bytes, 0x14) as usize;
let height = be32(bytes, 0x18) as usize;
if !(1..=4096).contains(&width) || !(1..=4096).contains(&height) {
return None;
}
let header = header_size(be32(bytes, 0x1c))?;
let n = (width as usize) * (height as usize) * 4;
if bytes.len() < header + n {
return None; // DXT/palettized/short variant — defer, don't misdecode
let tiles = be32(bytes, 0x1c) as usize;
let cols = width.div_ceil(TILE);
let rows = height.div_ceil(TILE);
// The field at 0x1c must be the tile count; otherwise it's a variant we don't
// decode (e.g. DXT / palettized) — defer rather than misdecode.
if tiles == 0 || tiles != cols * rows {
return None;
}
const TABLE: usize = 0x2c;
if bytes.len() < TABLE + tiles * 4 {
return None;
}
// A8R8G8B8 → RGBA8.
let src = &bytes[header..header + n];
let mut rgba = vec![0u8; n];
for (px, out) in src.chunks_exact(4).zip(rgba.chunks_exact_mut(4)) {
let (a, r, g, b) = (px[0], px[1], px[2], px[3]);
out[0] = r;
out[1] = g;
out[2] = b;
out[3] = a;
let mut rgba = vec![0u8; width * height * 4];
for ty in 0..rows {
for tx in 0..cols {
let tile = ty * cols + tx;
let pixels = be32(bytes, TABLE + tile * 4) as usize + TILE_HDR;
let tw = TILE.min(width - tx * TILE);
let th = TILE.min(height - ty * TILE);
if pixels + tw * th * 4 > bytes.len() {
return None; // truncated / not the layout we expect
}
for row in 0..th {
let mut s = pixels + row * tw * 4;
let mut d = ((ty * TILE + row) * width + tx * TILE) * 4;
for _ in 0..tw {
// A8R8G8B8 → RGBA8.
rgba[d] = bytes[s + 1];
rgba[d + 1] = bytes[s + 2];
rgba[d + 2] = bytes[s + 3];
rgba[d + 3] = bytes[s];
s += 4;
d += 4;
}
}
}
}
Some(T8adImage {
width,
height,
width: width as u32,
height: height as u32,
rgba,
})
}
@@ -95,13 +113,17 @@ pub fn parse(bytes: &[u8]) -> Option<T8adImage> {
mod tests {
use super::*;
/// Build a synthetic type-1 (header 64) T8aD with a known A8R8G8B8 pattern.
/// Build a synthetic single-tile T8aD (`w,h ≤ 256`) with a known A8R8G8B8
/// pattern: base header + 1-entry offset table + 16-byte tile header + pixels.
fn synth(w: u32, h: u32) -> Vec<u8> {
let mut b = vec![0u8; 64];
assert!(w <= 256 && h <= 256);
let mut b = vec![0u8; 0x2c];
b[0..4].copy_from_slice(&T8AD_MAGIC);
b[0x14..0x18].copy_from_slice(&w.to_be_bytes());
b[0x18..0x1c].copy_from_slice(&h.to_be_bytes());
b[0x1c..0x20].copy_from_slice(&1u32.to_be_bytes()); // type 1 → header 64
b[0x1c..0x20].copy_from_slice(&1u32.to_be_bytes()); // 1 tile
b.extend_from_slice(&0x30u32.to_be_bytes()); // offset table: tile 0 @ 0x30
b.extend_from_slice(&[0u8; 16]); // 16-byte tile header → pixels at 0x40
for i in 0..(w * h) {
b.extend_from_slice(&[(i & 0xff) as u8, 0x24, 0x63, 0xB2]); // A, R, G, B
}
@@ -121,8 +143,31 @@ mod tests {
}
#[test]
fn rejects_unknown_variant_and_short() {
// type field 7 (unknown) → None
fn assembles_row_major_tiles_via_offset_table() {
// 300×1 → 2 tiles: (0,0)=256×1 red, (1,0)=44×1 blue, each +16-byte header.
let (w, h): (u32, u32) = (300, 1);
let mut b = vec![0u8; 0x2c];
b[0..4].copy_from_slice(&T8AD_MAGIC);
b[0x14..0x18].copy_from_slice(&w.to_be_bytes());
b[0x18..0x1c].copy_from_slice(&h.to_be_bytes());
b[0x1c..0x20].copy_from_slice(&2u32.to_be_bytes()); // 2 tiles
let off0 = 0x2c + 2 * 4; // after the 2-entry table
let off1 = off0 + 16 + 256 * 4; // tile-0 header + its 256 pixels
b.extend_from_slice(&(off0 as u32).to_be_bytes());
b.extend_from_slice(&(off1 as u32).to_be_bytes());
b.extend_from_slice(&[0u8; 16]);
b.extend_from_slice(&[0xFF, 0xFF, 0, 0].repeat(256)); // A,R,G,B red
b.extend_from_slice(&[0u8; 16]);
b.extend_from_slice(&[0xFF, 0, 0, 0xFF].repeat(44)); // A,R,G,B blue
let img = parse(&b).expect("decodes");
assert_eq!((img.width, img.height), (300, 1));
assert_eq!(&img.rgba[0..4], &[0xFF, 0, 0, 0xFF]); // tile 0 → red
assert_eq!(&img.rgba[256 * 4..256 * 4 + 4], &[0, 0, 0xFF, 0xFF]); // tile 1 → blue
}
#[test]
fn rejects_wrong_tilecount_and_short() {
// tile-count field that isn't ceil(w/256)*ceil(h/256) → None
let mut b = synth(2, 2);
b[0x1c..0x20].copy_from_slice(&7u32.to_be_bytes());
assert!(parse(&b).is_none());

View File

@@ -221,6 +221,45 @@ impl X360Texture {
/// 3. Read its GPUTEXTURE_FETCH_CONSTANT (GPUFC) at descriptor +0x18
/// 4. De-tile the pixel data (if tiled) → linear layout
pub fn from_xpr2(bytes: &[u8]) -> Result<Self, TextureError> {
// XPR files are frequently PACKS of many textures; XPR_RES_INDEX picks
// the Nth texture resource (default 0) for RE/browse validation.
let want = std::env::var("XPR_RES_INDEX")
.ok()
.and_then(|v| v.trim().parse::<usize>().ok())
.unwrap_or(0);
Self::from_xpr2_index(bytes, want)
}
/// List the names of the texture resources (`TX2D` / `TXCM`) in an XPR2
/// container, in directory order — the same order [`from_xpr2_index`]
/// selects by. Non-texture resources (e.g. `XBG7`) are skipped.
pub fn texture_names(bytes: &[u8]) -> Vec<String> {
use std::io::Cursor;
let mut cur = Cursor::new(bytes);
let Ok(header) = Xpr2Header::read(&mut cur) else {
return Vec::new();
};
let mut names = Vec::new();
for _ in 0..header.num_resources {
let Ok(e) = Xpr2ResourceEntry::read(&mut cur) else {
break;
};
if e.is_texture() || e.is_cubemap() {
const DIR_BASE: usize = 0x10;
let no = e.name_offset as usize + DIR_BASE;
let name = bytes
.get(no..)
.and_then(|s| s.iter().position(|&c| c == 0).map(|p| &s[..p]))
.map(|s| String::from_utf8_lossy(s).into_owned())
.unwrap_or_default();
names.push(name);
}
}
names
}
/// Decode the `want`-th texture resource (`TX2D` / `TXCM`, directory order).
pub fn from_xpr2_index(bytes: &[u8], want: usize) -> Result<Self, TextureError> {
use std::io::Cursor;
let mut cur = Cursor::new(bytes);
@@ -236,12 +275,6 @@ impl X360Texture {
// Select a texture resource. TX2D = 2D texture; TXCM = cubemap
// (skybox / backdrop) — same 52-byte descriptor + GPUFC layout, but the
// pixel section holds 6 faces. For a preview we decode face 0.
// XPR files are frequently PACKS of many textures; XPR_RES_INDEX picks
// the Nth texture resource (default 0) for RE/browse validation.
let want = std::env::var("XPR_RES_INDEX")
.ok()
.and_then(|v| v.trim().parse::<usize>().ok())
.unwrap_or(0);
let tex_entry = entries.iter()
.filter(|e| e.is_texture() || e.is_cubemap())
.nth(want)

View File

@@ -0,0 +1,426 @@
//! Integration test: decode XBG7 geometry from REAL `.xpr` model files.
//!
//! Uses loose files extracted from the retail disc (models live in
//! `hidden/resource3d/*.xpr`). Skipped unless the directory is found; point
//! `SYLPHEED_RES3D` at it to override.
//!
//! Run: `cargo test -p sylpheed-formats --test mesh_disc -- --ignored --nocapture`
use std::path::PathBuf;
use sylpheed_formats::mesh::{material_groups, node_transforms, submesh_albedos, Xbg7Model};
fn res3d_dir() -> Option<PathBuf> {
if let Ok(p) = std::env::var("SYLPHEED_RES3D") {
let p = PathBuf::from(p);
if p.is_dir() {
return Some(p);
}
}
let default =
PathBuf::from("/home/fabi/RE Project Sylpheed/sylph_extract/hidden/resource3d");
default.is_dir().then_some(default)
}
#[test]
#[ignore = "requires extracted disc models — set SYLPHEED_RES3D"]
fn hero_ship_submesh_material_graph() {
let Some(dir) = res3d_dir() else {
eprintln!("SKIP: resource3d dir not found (set SYLPHEED_RES3D)");
return;
};
// The XBG7 node/material graph tags each sub-mesh record with its own albedo
// (`rou_…_col` → TX2D name). DeltaSaber_T `f001`: the 7 detail parts use
// bdy_04/06/07 — NOT the body's bdy_01a — which is the per-part texturing the
// viewer needs so fins aren't painted with the hull map.
let bytes = std::fs::read(dir.join("DeltaSaber_T.xpr")).unwrap();
let mut map = std::collections::HashMap::new();
for (v, i, name) in submesh_albedos(&bytes, "f001") {
map.insert((v, i), name);
}
// (vtx, idx) → expected albedo (rou_ stripped = TX2D name).
assert_eq!(map.get(&(314, 645)).map(String::as_str), Some("f001_bdy_04_col"));
assert_eq!(map.get(&(48, 84)).map(String::as_str), Some("f001_bdy_04_col"));
assert_eq!(map.get(&(96, 180)).map(String::as_str), Some("f001_bdy_06_col"));
assert_eq!(map.get(&(60, 108)).map(String::as_str), Some("f001_bdy_07_col"));
// Names strip cleanly to real TX2D resources.
let tex = sylpheed_formats::texture::X360Texture::texture_names(&bytes);
for name in map.values() {
assert!(tex.iter().any(|t| t == name), "albedo {name} not a TX2D resource");
}
}
#[test]
#[ignore = "requires extracted disc models — set SYLPHEED_RES3D"]
fn hero_ship_body_splits_by_material() {
let Some(dir) = res3d_dir() else {
eprintln!("SKIP: resource3d dir not found (set SYLPHEED_RES3D)");
return;
};
// The merged body (sub0 = 24561 indices / 8187 tris) is drawn as 9 groups
// sharing one index buffer, each with its own albedo (bdy_01a hull, bdy_01b,
// bdy_02/03, and the `daiza` hangar stand). `material_groups` recovers the
// cumulative-offset chain so the viewer can texture each slice.
let bytes = std::fs::read(dir.join("DeltaSaber_T.xpr")).unwrap();
let groups = material_groups(&bytes, "f001", 24561);
assert_eq!(groups.len(), 9, "body must split into 9 material groups");
// Offsets chain from 0 and cover the whole index buffer exactly.
let mut cursor = 0usize;
for g in &groups {
assert_eq!(g.idx_offset, cursor, "group offsets must be contiguous");
cursor += g.idx_count;
}
assert_eq!(cursor, 24561, "groups must cover all body indices");
// Expected per-group albedo (verified against the GPU draw log).
let albedos: Vec<&str> = groups.iter().map(|g| g.albedo.as_str()).collect();
assert_eq!(
albedos,
[
"f001_bdy_01a_col",
"f001_bdy_01a_col",
"f001_bdy_01a_col",
"f001_bdy_01b_col",
"f001_bdy_01a_col",
"f001_bdy_01b_col",
"f001_bdy_02_col",
"f001_bdy_03_col",
"f001_bdy_daiza_col",
]
);
// A single-material detail part (sub1 = 645 indices) → one group, bdy_04.
let fin = material_groups(&bytes, "f001", 645);
assert_eq!(fin.len(), 1);
assert_eq!(fin[0].albedo, "f001_bdy_04_col");
assert_eq!(fin[0].idx_offset, 0);
}
#[test]
#[ignore = "requires extracted disc models — set SYLPHEED_RES3D"]
fn hero_ship_node_transforms_move_fins_to_tail() {
let Some(dir) = res3d_dir() else {
eprintln!("SKIP: resource3d dir not found (set SYLPHEED_RES3D)");
return;
};
// For the DeltaSaber, node_transforms returns the MEASURED runtime placements
// (Canary draw-log WVP; see mesh.rs::DELTASABER_MEASURED) — the fin assemblies
// are mounted OUT on the nacelles (world X ≈ 4.4..6.6), an offset the static
// graph omits. Body identity; each part a left/right mirrored pair. Vertex
// space: X right, Y up, Z fore/aft (engines at Z).
let bytes = std::fs::read(dir.join("DeltaSaber_T.xpr")).unwrap();
let places = node_transforms(&bytes, "f001");
assert!(!places.is_empty(), "expected node placements");
// Body (sub 0) identity, drawn once.
let body: Vec<_> = places.iter().filter(|p| p.sub_index == 0).collect();
assert_eq!(body.len(), 1, "body drawn once");
assert!(body[0].t.iter().all(|c| c.abs() < 0.1), "body must not translate");
assert!(!body[0].reflect, "body is not mirrored");
// Fin bdy_04 (sub 1): mirrored pair near the tail, mounted OUT on the nacelle
// (|X| ≈ 5.2, not the graph's inboard ≈1.1).
let fins: Vec<_> = places.iter().filter(|p| p.sub_index == 1).collect();
assert_eq!(fins.len(), 2, "V-tail is a mirrored pair");
assert!(fins.iter().all(|f| (f.t[2] + 9.72).abs() < 0.3), "fin fore/aft ≈ 9.7");
assert!(fins.iter().all(|f| f.t[0].abs() > 4.0), "fin mounted out on the nacelle");
assert!(fins[0].t[0] * fins[1].t[0] < 0.0, "the pair mirrors across X");
assert!(fins.iter().any(|f| f.reflect), "one of the pair is reflected");
// Winglet bdy_06 (sub 4): on the nacelle (|X| ≈ 6.6, Z ≈ 15.5).
let wings: Vec<_> = places.iter().filter(|p| p.sub_index == 4).collect();
assert_eq!(wings.len(), 2, "L/R winglet pair");
assert!(wings.iter().all(|p| p.t[0].abs() > 6.0 && (p.t[2] + 15.5).abs() < 0.3),
"winglets out on the nacelle");
assert!(wings[0].t[0] * wings[1].t[0] < 0.0, "winglets mirror across X");
// Small fin bdy_10 (sub 6): inboard nacelle stub (|X| ≈ 4.45, Z ≈ 17).
let sfins: Vec<_> = places.iter().filter(|p| p.sub_index == 6).collect();
assert_eq!(sfins.len(), 2, "L/R small-fin pair");
assert!(sfins.iter().all(|p| (p.t[0].abs() - 4.45).abs() < 0.3), "small fins on the stub");
assert!(sfins[0].t[0] * sfins[1].t[0] < 0.0, "small fins mirror across X");
}
#[test]
#[ignore = "requires extracted disc models — set SYLPHEED_RES3D"]
fn weapon_model_decodes_to_expected_geometry() {
let Some(dir) = res3d_dir() else {
eprintln!("SKIP: resource3d dir not found (set SYLPHEED_RES3D)");
return;
};
// rou_f001_wep_00 = the player ship's first weapon: 1 sub-mesh,
// 215 vertices, 364 triangles (verified by hex analysis).
let bytes = std::fs::read(dir.join("rou_f001_wep_00.xpr")).unwrap();
let model = Xbg7Model::from_xpr2(&bytes).expect("weapon must decode");
assert_eq!(model.meshes.len(), 1);
let (v, t) = model.totals();
assert_eq!(v, 215, "vertex count");
assert_eq!(t, 364, "triangle count");
let m = &model.meshes[0];
assert_eq!(m.positions.len(), 215);
assert_eq!(m.uvs.len(), 215);
assert_eq!(m.normals.len(), 215);
assert_eq!(m.indices.len(), 1092);
// every index in range
assert!(m.indices.iter().all(|&i| (i as usize) < m.positions.len()));
// positions are real geometry within the model's ~2-unit bbox
let ys: Vec<f32> = m.positions.iter().map(|p| p[1]).collect();
let span = ys.iter().cloned().fold(f32::MIN, f32::max)
- ys.iter().cloned().fold(f32::MAX, f32::min);
assert!(span > 1.0 && span < 10.0, "y-span {span} out of range");
// Correct vertex alignment ⇒ normals are unit-length (the pin for the
// +12 vertex offset) and UVs land in a sane texture range.
let mean_nlen: f32 = m
.normals
.iter()
.map(|n| (n[0] * n[0] + n[1] * n[1] + n[2] * n[2]).sqrt())
.sum::<f32>()
/ m.normals.len() as f32;
assert!((mean_nlen - 1.0).abs() < 0.05, "mean |normal| {mean_nlen} ≠ 1");
assert!(
m.uvs.iter().all(|uv| uv[0] > -0.1 && uv[0] < 2.0 && uv[1] > -0.1 && uv[1] < 2.0),
"UVs out of expected [0,1]-ish range"
);
}
#[test]
#[ignore = "requires extracted disc models — set SYLPHEED_RES3D"]
fn declaration_driven_decode_covers_expected_model_count() {
let Some(dir) = res3d_dir() else {
return;
};
let mut ok = 0usize;
let mut total = 0usize;
for entry in std::fs::read_dir(&dir).unwrap() {
let path = entry.unwrap().path();
if path.extension().and_then(|e| e.to_str()) != Some("xpr") {
continue;
}
total += 1;
let bytes = std::fs::read(&path).unwrap();
if let Ok(model) = Xbg7Model::from_xpr2(&bytes) {
if !model.meshes.is_empty() {
// Every decoded model must be self-consistent (indices in range,
// and unit normals where present).
for m in &model.meshes {
assert!(
m.indices.iter().all(|&i| (i as usize) < m.positions.len()),
"{path:?}: index out of range"
);
}
ok += 1;
}
}
}
eprintln!("XBG7 declaration-driven decode: {ok}/{total} models");
// Variable-stride declaration parsing lifted coverage vs the old fixed
// stride-24 decoder (25 → 36 fully-validated models; multi-sub-mesh models
// whose later sub-mesh offset isn't yet handled are still declined whole).
assert!(ok >= 35, "coverage regressed: only {ok}/{total} decoded");
}
#[test]
#[ignore = "requires extracted disc models — set SYLPHEED_RES3D"]
fn complex_body_mesh_is_declined_not_garbage() {
let Some(dir) = res3d_dir() else {
return;
};
// The hero-ship body uses the multi-stream layout we do not decode; it must
// be cleanly rejected, never returned as partial geometry.
let bytes = std::fs::read(dir.join("DeltaSaber_A.xpr")).unwrap();
match Xbg7Model::from_xpr2(&bytes) {
Err(_) => {} // expected: declined
Ok(m) => {
// If it ever does decode, it must at least be self-consistent.
for mesh in &m.meshes {
assert!(mesh.indices.iter().all(|&i| (i as usize) < mesh.positions.len()));
}
}
}
}
/// The hero ship's **grouped-pool** layout decodes fully: `DeltaSaber_T.xpr`'s
/// `f001` resource is one vertex+index pool shared by 8 sub-meshes (body + 7
/// detail parts). Reversed statically and cross-checked against a Canary GPU
/// draw-log capture — every sub-mesh's stored normals agree with its triangle
/// winding (0 degenerate, full vertex coverage). This is the layout the old
/// per-block adjacency anchor rendered as a spiky phantom.
#[test]
#[ignore = "requires extracted disc models — set SYLPHEED_RES3D"]
fn hero_ship_grouped_pool_decodes() {
let Some(dir) = res3d_dir() else {
eprintln!("SKIP: resource3d dir not found (set SYLPHEED_RES3D)");
return;
};
let bytes = std::fs::read(dir.join("DeltaSaber_T.xpr")).unwrap();
let models = Xbg7Model::stage_models(&bytes);
// The neutral pose `f001` (the mnv*/turn180 resources are animation poses).
let f001 = models.iter().find(|m| m.name == "f001").expect("f001 decoded");
assert_eq!(f001.meshes.len(), 8, "body + 7 detail sub-meshes");
let (v, t) = f001.totals();
assert_eq!(v, 11607, "vertex total across sub-meshes");
assert_eq!(t, 8650, "triangle total (body 8187 + 7 parts)");
// Sub-mesh 0 is the body: exactly the draw-log-verified geometry.
let body = &f001.meshes[0];
assert_eq!(body.positions.len(), 10891);
assert_eq!(body.indices.len(), 24561);
for (i, m) in f001.meshes.iter().enumerate() {
// Every index in range.
assert!(
m.indices.iter().all(|&ix| (ix as usize) < m.positions.len()),
"sub{i}: index out of range"
);
// Correct alignment ⇒ unit normals, and the winding agrees with them
// (the decisive correctness signal, ~1.0, not the ~0.5 of a mis-carve).
let n = m.normals.len();
assert_eq!(n, m.positions.len(), "sub{i}: a normal per vertex");
let mean_nlen: f32 = m
.normals
.iter()
.map(|nv| (nv[0] * nv[0] + nv[1] * nv[1] + nv[2] * nv[2]).sqrt())
.sum::<f32>()
/ n as f32;
assert!((mean_nlen - 1.0).abs() < 0.05, "sub{i}: mean |normal| {mean_nlen} ≠ 1");
let mut agree = 0usize;
let mut counted = 0usize;
for tri in m.indices.chunks_exact(3) {
let (a, b, c) = (tri[0] as usize, tri[1] as usize, tri[2] as usize);
let (pa, pb, pc) = (m.positions[a], m.positions[b], m.positions[c]);
let u = [pb[0] - pa[0], pb[1] - pa[1], pb[2] - pa[2]];
let w = [pc[0] - pa[0], pc[1] - pa[1], pc[2] - pa[2]];
let f = [
u[1] * w[2] - u[2] * w[1],
u[2] * w[0] - u[0] * w[2],
u[0] * w[1] - u[1] * w[0],
];
if f[0] * f[0] + f[1] * f[1] + f[2] * f[2] < 1e-12 {
continue;
}
let sn = [
m.normals[a][0] + m.normals[b][0] + m.normals[c][0],
m.normals[a][1] + m.normals[b][1] + m.normals[c][1],
m.normals[a][2] + m.normals[b][2] + m.normals[c][2],
];
if f[0] * sn[0] + f[1] * sn[1] + f[2] * sn[2] > 0.0 {
agree += 1;
}
counted += 1;
}
let na = agree as f32 / counted.max(1) as f32;
assert!(
na.max(1.0 - na) > 0.95,
"sub{i}: winding-vs-normal agreement {na} — mis-carved (should be ~1.0 or ~0.0)"
);
}
}
/// Stage containers decode multiple enemy/prop sub-models via content anchoring.
/// Prints coverage; asserts the known-good Stage_S10 meshes decode with sane geometry.
#[test]
#[ignore]
fn stage_models_decode() {
use sylpheed_formats::mesh::Xbg7Model;
let dir = std::env::var("SYLPHEED_RES3D").unwrap_or_else(|_| {
"/home/fabi/RE Project Sylpheed/sylph_extract/hidden/resource3d".to_string()
});
let path = format!("{dir}/Stage_S10.xpr");
let bytes = std::fs::read(&path).expect("read Stage_S10");
let models = Xbg7Model::stage_models(&bytes);
for m in &models {
let (v, t) = m.totals();
let mut lo = [f32::MAX; 3];
let mut hi = [f32::MIN; 3];
for sub in &m.meshes {
for p in &sub.positions {
for a in 0..3 {
lo[a] = lo[a].min(p[a]);
hi[a] = hi[a].max(p[a]);
}
}
}
let ext = [hi[0] - lo[0], hi[1] - lo[1], hi[2] - lo[2]];
println!(
" {:16} v={v} t={t} bbox=[{:.1},{:.1},{:.1}]",
m.name, ext[0], ext[1], ext[2]
);
}
// Known: e003 (main enemy, ~1400 tris) must be among the decoded models.
let e003 = models.iter().find(|m| m.name == "e003").expect("e003 decoded");
let (v, t) = e003.totals();
assert_eq!(v, 2383, "e003 vertex count");
assert!(t > 1400, "e003 triangle count {t}");
assert!(models.len() >= 3, "at least 3 stage sub-models, got {}", models.len());
}
/// Timing/coverage sweep across all stage containers (manual; release recommended).
#[test]
#[ignore]
fn stage_models_sweep() {
use sylpheed_formats::mesh::Xbg7Model;
use std::time::Instant;
let dir = std::env::var("SYLPHEED_RES3D").unwrap_or_else(|_| {
"/home/fabi/RE Project Sylpheed/sylph_extract/hidden/resource3d".to_string()
});
let mut names: Vec<_> = std::fs::read_dir(&dir)
.unwrap()
.filter_map(|e| e.ok().map(|e| e.file_name().into_string().unwrap()))
.filter(|n| n.starts_with("Stage_") && n.ends_with(".xpr"))
.collect();
names.sort();
let mut tot = 0usize;
for n in &names {
let bytes = std::fs::read(format!("{dir}/{n}")).unwrap();
let t0 = Instant::now();
let models = Xbg7Model::stage_models(&bytes);
let dt = t0.elapsed().as_millis();
tot += models.len();
println!("{n:16} {:>4} MB {:>3} models {:>5} ms", bytes.len()/1_000_000, models.len(), dt);
}
println!("TOTAL stage sub-models decoded: {tot}");
}
/// Quality audit: for a big stage, verify decoded models are real geometry
/// (bounded bbox, low full-mesh degeneracy) rather than false anchors.
#[test]
#[ignore]
fn stage_models_quality_audit() {
use sylpheed_formats::mesh::Xbg7Model;
let dir = std::env::var("SYLPHEED_RES3D").unwrap_or_else(|_| {
"/home/fabi/RE Project Sylpheed/sylph_extract/hidden/resource3d".to_string()
});
let bytes = std::fs::read(format!("{dir}/Stage_S07.xpr")).unwrap();
let models = Xbg7Model::stage_models(&bytes);
let (mut small, mut mid, mut huge, mut dupnames) = (0, 0, 0, 0);
let mut seen = std::collections::HashSet::new();
let mut worst_deg = 0.0f32;
for m in &models {
if !seen.insert(m.name.clone()) { dupnames += 1; }
let mut lo = [f32::MAX; 3]; let mut hi = [f32::MIN; 3];
let mut deg = 0usize; let mut tot = 0usize;
for sub in &m.meshes {
for p in &sub.positions { for a in 0..3 { lo[a]=lo[a].min(p[a]); hi[a]=hi[a].max(p[a]); } }
for tri in sub.indices.chunks_exact(3) {
let (a,b,c)=(tri[0] as usize,tri[1] as usize,tri[2] as usize);
if a>=sub.positions.len()||b>=sub.positions.len()||c>=sub.positions.len(){continue;}
let pa=sub.positions[a]; let pb=sub.positions[b]; let pc=sub.positions[c];
let u=[pb[0]-pa[0],pb[1]-pa[1],pb[2]-pa[2]];
let v=[pc[0]-pa[0],pc[1]-pa[1],pc[2]-pa[2]];
let cx=[u[1]*v[2]-u[2]*v[1],u[2]*v[0]-u[0]*v[2],u[0]*v[1]-u[1]*v[0]];
if 0.5*(cx[0]*cx[0]+cx[1]*cx[1]+cx[2]*cx[2]).sqrt()<1e-9 { deg+=1; }
tot+=1;
}
}
let ext = (hi[0]-lo[0]).max(hi[1]-lo[1]).max(hi[2]-lo[2]);
let df = if tot>0 { deg as f32/tot as f32 } else {1.0};
worst_deg = worst_deg.max(df);
if ext < 200.0 { small += 1; } else if ext < 5000.0 { mid += 1; } else { huge += 1; }
}
println!("S07: {} models | small(<200u)={} mid={} huge(>5k)={} | dup-names={} | worst full-mesh degeneracy={:.1}%",
models.len(), small, mid, huge, dupnames, worst_deg*100.0);
// Each XBG7 resource must anchor to a *distinct* block (no collisions), the
// bulk must be bounded-scale geometry, and none may be mostly-degenerate.
assert_eq!(dupnames, 0, "no two resources should anchor to the same block");
assert!(small + mid > models.len() * 9 / 10, "≥90% bounded-scale geometry");
assert!(huge < models.len() / 20, "few huge (skybox-plane) models");
assert!(worst_deg < 0.35, "no model should be mostly-degenerate");
}

View File

@@ -0,0 +1,102 @@
//! Real-disc test for the movie manifest → voice binding. Skipped without
//! `SYLPHEED_DISC`.
use std::path::PathBuf;
use sylpheed_formats::movie_manifest;
use sylpheed_formats::slb::VoiceLang;
use sylpheed_formats::PakArchive;
fn disc_root() -> Option<PathBuf> {
let p = PathBuf::from(std::env::var("SYLPHEED_DISC").ok()?);
p.join("dat").is_dir().then_some(p)
}
/// Read the manifest + `eng\sounds.tbl` out of `tables.pak`.
fn load_manifest_and_sounds(root: &PathBuf) -> (Vec<u8>, Vec<u8>) {
let pak = PakArchive::open(root.join("dat/tables.pak")).unwrap();
let manifest = pak
.entries()
.iter()
.find_map(|e| pak.read(e).ok().filter(|b| movie_manifest::is_manifest(b)))
.expect("movie manifest present in tables.pak");
let sounds = pak
.read_by_name("eng\\sounds.tbl")
.expect("sounds.tbl present")
.unwrap();
(manifest, sounds)
}
#[test]
fn binds_and_resolves_movie_voice() {
let Some(root) = disc_root() else {
eprintln!("SKIP: set SYLPHEED_DISC");
return;
};
let (manifest, sounds) = load_manifest_and_sounds(&root);
let entries = movie_manifest::parse(&manifest);
assert!(entries.len() > 90, "expected ~101 movies, got {}", entries.len());
// Standard story movie → VOICE_<movie> in <lang>\Movie\.
assert_eq!(
movie_manifest::voice_token(&manifest, "S13A").as_deref(),
Some("VOICE_S13A")
);
assert_eq!(
movie_manifest::resolve_voice_entry(&manifest, &sounds, "S13A", VoiceLang::English)
.as_deref(),
Some("eng\\Movie\\VOICE_S13A.slb")
);
// A resupply movie bound to an in-mission radio clip → VOICE_D_* in \etc\.
// (This is what a `VOICE_<movie>` guess would miss entirely.)
assert_eq!(
movie_manifest::voice_token(&manifest, "hokyu_LS_s02A").as_deref(),
Some("VOICE_D_450")
);
assert_eq!(
movie_manifest::resolve_voice_entry(
&manifest,
&sounds,
"hokyu_LS_s02A",
VoiceLang::English
)
.as_deref(),
Some("eng\\etc\\VOICE_D_450.slb")
);
// A resupply movie the game leaves unbound in the manifest → no DIRECT
// binding. (Extending to unbound movies by shared demo line was verified
// WRONG against the running game, so we do NOT resolve these.)
let e = entries
.iter()
.find(|e| e.movie == "hokyu_DS_s13A")
.expect("hokyu_DS_s13A present in manifest");
assert_eq!(e.voice_token, None, "hokyu_DS_s13A has no direct voice binding");
assert_eq!(
movie_manifest::resolve_voice_entry(&manifest, &sounds, "hokyu_DS_s13A", VoiceLang::English),
None
);
// Every resolved voice entry must actually exist in sound.pak.
let snd = std::fs::read(root.join("dat/sound.pak")).unwrap();
let keys: std::collections::HashSet<u32> = PakArchive::parse_toc(&snd)
.unwrap()
.iter()
.map(|e| e.name_hash)
.collect();
let mut resolved = 0;
let mut missing = Vec::new();
for e in &entries {
if let Some(full) =
movie_manifest::resolve_voice_entry(&manifest, &sounds, &e.movie, VoiceLang::English)
{
resolved += 1;
if !keys.contains(&sylpheed_formats::hash::name_hash(&full)) {
missing.push(full);
}
}
}
assert!(resolved > 80, "expected 80+ voiced movies, got {resolved}");
assert!(missing.is_empty(), "resolved but absent in sound.pak: {missing:?}");
}

View File

@@ -0,0 +1,63 @@
//! Real-disc test for the movie subtitle chain. Skipped when the extracted disc
//! is absent (set `SYLPHEED_DISC` to the extract root to enable).
use std::path::PathBuf;
use sylpheed_formats::movie_subtitle::{self, SubLang};
use sylpheed_formats::PakArchive;
fn disc_root() -> Option<PathBuf> {
let p = PathBuf::from(std::env::var("SYLPHEED_DISC").ok()?);
p.join("dat").is_dir().then_some(p)
}
#[test]
fn resolves_english_radio_subtitles() {
let Some(root) = disc_root() else {
eprintln!("SKIP: set SYLPHEED_DISC");
return;
};
let lang = SubLang::English;
let lang_pak = PakArchive::open(root.join(format!("dat/movie/{}.pak", lang.pak_code()))).unwrap();
let text_pak =
PakArchive::open(root.join(format!("dat/GP_MAIN_GAME_{}.pak", lang.game_code()))).unwrap();
// A radio movie: real timed English dialogue.
let cues = movie_subtitle::load("RT01C_1", &lang_pak, &text_pak);
assert!(!cues.is_empty(), "RT01C_1 should have timed cues");
assert!(cues[0].start >= 0.0 && cues[0].start < 5.0);
assert!(
cues.iter().any(|c| c.text.contains("Rhino Leader")),
"expected recognizable dialogue, got: {:?}",
cues.iter().map(|c| &c.text).collect::<Vec<_>>()
);
// No MSG_DEMO keys should leak into resolved caption text.
assert!(
cues.iter().all(|c| !c.text.contains("MSG_DEMO")),
"unresolved key leaked: {:?}",
cues.iter().map(|c| &c.text).collect::<Vec<_>>()
);
// A resupply movie: the canonical resupply line.
let hokyu = movie_subtitle::load("hokyu_DS_s02A", &lang_pak, &text_pak);
assert!(hokyu.iter().any(|c| c.text.contains("Resupply complete")));
// S00A carries inline narration text (not MSG_DEMO refs) with timecodes.
let story = movie_subtitle::load("S00A", &lang_pak, &text_pak);
assert!(!story.is_empty() && story.iter().any(|c| c.text.contains("27th century")));
// S13A stores a two-line caption as two consecutive text tokens sharing one
// timing; both lines must survive (the "Look at it father" regression — the
// first line used to be dropped).
let s13 = movie_subtitle::load("S13A", &lang_pak, &text_pak);
let multiline = s13
.iter()
.find(|c| c.text.contains("Look at it father"))
.expect("S13A should contain the 'Look at it father' caption");
assert_eq!(
multiline.text, "Look at it father\n& beautiful isn't it",
"both lines of the caption must be present"
);
assert!((multiline.start - 74.8).abs() < 0.1, "start {}", multiline.start);
}

View File

@@ -0,0 +1,104 @@
//! Real-disc test: read a movie voice `.slb` out of the segmented `sound.pak`
//! and rebuild a decodable XMA1 RIFF. Skipped without `SYLPHEED_DISC`.
use std::fs::File;
use std::io::{Read, Seek, SeekFrom};
use std::path::PathBuf;
use sylpheed_formats::hash::name_hash;
use sylpheed_formats::slb::{self, VoiceLang};
use sylpheed_formats::PakArchive;
fn disc_root() -> Option<PathBuf> {
let p = PathBuf::from(std::env::var("SYLPHEED_DISC").ok()?);
p.join("dat").is_dir().then_some(p)
}
/// Read `[off, off+size)` from `dat/sound.p00..` (segments concatenated).
fn read_range(root: &PathBuf, mut off: u64, size: usize) -> Vec<u8> {
let mut out = Vec::with_capacity(size);
let mut need = size;
for i in 0..100u32 {
if need == 0 {
break;
}
let seg = root.join(format!("dat/sound.p{i:02}"));
let Ok(meta) = std::fs::metadata(&seg) else { break };
let len = meta.len();
if off >= len {
off -= len;
continue;
}
let mut f = File::open(&seg).unwrap();
f.seek(SeekFrom::Start(off)).unwrap();
let take = need.min((len - off) as usize);
let start = out.len();
out.resize(start + take, 0);
f.read_exact(&mut out[start..]).unwrap();
need -= take;
off = 0;
}
out
}
#[test]
fn reads_and_rebuilds_movie_voice_riff() {
let Some(root) = disc_root() else {
eprintln!("SKIP: set SYLPHEED_DISC");
return;
};
let toc = std::fs::read(root.join("dat/sound.pak")).unwrap();
let entries = PakArchive::parse_toc(&toc).unwrap();
assert_eq!(entries.len(), 9519);
for movie in ["S00A", "RT07A", "RT02D_1"] {
// RT02D_1 is a headerless variant; S00A/RT07A embed a RIFF.
let key = name_hash(&slb::movie_voice_name(movie, VoiceLang::English));
let idx = entries
.binary_search_by_key(&key, |e| e.name_hash)
.unwrap_or_else(|_| panic!("{movie} voice missing"));
let e = &entries[idx];
let bank = read_range(&root, e.offset as u64, e.comp_size as usize);
let riff = slb::to_xma_riff(&bank).expect("rebuild RIFF");
assert_eq!(&riff[0..4], b"RIFF", "{movie}");
assert_eq!(&riff[8..12], b"WAVE", "{movie}");
// fmt chunk advertises XMA1 (tag 0x0165).
let fpos = riff.windows(4).position(|w| w == b"fmt ").unwrap();
let tag = u16::from_le_bytes([riff[fpos + 8], riff[fpos + 9]]);
assert_eq!(tag, 0x0165, "{movie} should be XMA1");
// data chunk carries the XMA payload.
let dpos = riff.windows(4).position(|w| w == b"data").unwrap();
let dsz = u32::from_le_bytes([
riff[dpos + 4],
riff[dpos + 5],
riff[dpos + 6],
riff[dpos + 7],
]) as usize;
assert!(dsz > 10_000, "{movie} data too small: {dsz}");
assert_eq!(riff.len(), dpos + 8 + dsz, "{movie} data length consistent");
}
}
#[test]
fn enumerates_voice_clips_from_sounds_tbl() {
let Some(root) = disc_root() else {
eprintln!("SKIP: set SYLPHEED_DISC");
return;
};
let pak = PakArchive::open(root.join("dat/tables.pak")).unwrap();
let tbl = pak.read_by_name("eng\\sounds.tbl").unwrap().unwrap();
let clips = slb::list_voice_clips(&tbl, VoiceLang::English);
assert!(clips.len() > 1000, "expected many voice clips, got {}", clips.len());
// Every clip is an eng voice path present in sound.pak.
let sound = std::fs::read(root.join("dat/sound.pak")).unwrap();
let keys: std::collections::HashSet<u32> =
PakArchive::parse_toc(&sound).unwrap().iter().map(|e| e.name_hash).collect();
let present = clips.iter().filter(|c| keys.contains(&name_hash(&c.name))).count();
assert!(
present as f32 / clips.len() as f32 > 0.95,
"{present}/{} clips resolve in sound.pak",
clips.len()
);
assert!(clips.iter().any(|c| c.speaker == "ADAN"));
}

View File

@@ -8,6 +8,10 @@
use bevy::prelude::*;
use bevy::input::mouse::{MouseMotion, MouseWheel};
use bevy::render::render_asset::RenderAssetUsages;
use bevy::render::render_resource::{
Extent3d, TextureDimension, TextureFormat, TextureViewDescriptor, TextureViewDimension,
};
pub struct OrbitCameraPlugin;
@@ -32,6 +36,10 @@ pub struct OrbitCamera {
pub orbit_sensitivity: f32,
pub zoom_sensitivity: f32,
pub pan_sensitivity: f32,
/// Zoom limits — set when framing a model so large scenes can be pulled back
/// far enough (the old fixed 0.5..50 cap trapped the camera inside big stages).
pub min_radius: f32,
pub max_radius: f32,
}
impl Default for OrbitCamera {
@@ -44,29 +52,114 @@ impl Default for OrbitCamera {
orbit_sensitivity: 0.005,
zoom_sensitivity: 0.3,
pan_sensitivity: 0.003,
min_radius: 0.05,
max_radius: 500.0,
}
}
}
fn spawn_camera(mut commands: Commands) {
fn spawn_camera(mut commands: Commands, mut images: ResMut<Assets<Image>>) {
let orbit = OrbitCamera::default();
let transform = orbit_transform(&orbit);
// The game's ship shader reflects a shared HDR environment cubemap off the
// hull (three cube samples — see the shader RE). Reproduce that "look" with a
// procedural sky cube feeding Bevy's image-based lighting, so metallic
// surfaces reflect an environment instead of reading flat. Procedural (not a
// bundled KTX2) so it also works on WASM with no extra assets.
let env = make_env_cubemap(&mut images);
commands.spawn((
Camera3d::default(),
// Wide near/far so both tiny weapons and multi-thousand-unit stages fit;
// the planes are re-scaled to the zoom distance each frame (below).
Projection::Perspective(PerspectiveProjection {
near: 0.05,
far: 100_000.0,
..default()
}),
transform,
orbit,
EnvironmentMapLight {
diffuse_map: env.clone(),
specular_map: env,
intensity: 900.0,
..default()
},
));
}
/// Build a small procedural sky cubemap (6 faces) for image-based lighting.
///
/// A vertical gradient (bright zenith → warm horizon → dark ground) plus a warm
/// "sun" highlight roughly where the game's key light sits — enough for a
/// metallic hull to read as reflective rather than flat. Stored as an sRGB cube
/// (filterable, no half-float encoding needed); a single mip means reflections
/// stay sharp (fine for a shiny ship).
fn make_env_cubemap(images: &mut Assets<Image>) -> Handle<Image> {
let size: i32 = 64;
// Approx key-light direction from the RE (PS c32 ≈ (0,-0.76,-0.65), i.e. light
// arriving from top-front); place the bright spot there.
let sun = Vec3::new(0.0, 0.85, 0.5).normalize();
let zenith = Vec3::new(0.70, 0.80, 0.95);
let horizon = Vec3::new(0.42, 0.45, 0.50);
let ground = Vec3::new(0.09, 0.09, 0.11);
let srgb = |c: f32| (c.clamp(0.0, 1.0).powf(1.0 / 2.2) * 255.0).round() as u8;
let mut data: Vec<u8> = Vec::with_capacity((size * size * 6 * 4) as usize);
for face in 0..6 {
for y in 0..size {
for x in 0..size {
let u = (x as f32 + 0.5) / size as f32 * 2.0 - 1.0;
let v = (y as f32 + 0.5) / size as f32 * 2.0 - 1.0;
// Standard cube-face → direction mapping.
let d = match face {
0 => Vec3::new(1.0, -v, -u),
1 => Vec3::new(-1.0, -v, u),
2 => Vec3::new(u, 1.0, v),
3 => Vec3::new(u, -1.0, -v),
4 => Vec3::new(u, -v, 1.0),
_ => Vec3::new(-u, -v, -1.0),
}
.normalize();
let mut col = if d.y >= 0.0 {
horizon.lerp(zenith, (d.y).powf(0.6))
} else {
horizon.lerp(ground, (-d.y).powf(0.5))
};
let s = d.dot(sun).max(0.0).powf(60.0);
col += Vec3::new(1.0, 0.85, 0.6) * s * 1.5;
data.extend_from_slice(&[srgb(col.x), srgb(col.y), srgb(col.z), 255]);
}
}
}
let mut image = Image::new(
Extent3d {
width: size as u32,
height: size as u32,
depth_or_array_layers: 6,
},
TextureDimension::D2,
data,
TextureFormat::Rgba8UnormSrgb,
RenderAssetUsages::RENDER_WORLD,
);
image.texture_view_descriptor = Some(TextureViewDescriptor {
dimension: Some(TextureViewDimension::Cube),
..default()
});
images.add(image)
}
fn orbit_camera(
mut query: Query<(&mut OrbitCamera, &mut Transform)>,
mut query: Query<(&mut OrbitCamera, &mut Transform, &mut Projection)>,
mouse_buttons: Res<ButtonInput<MouseButton>>,
keys: Res<ButtonInput<KeyCode>>,
mut mouse_motion: EventReader<MouseMotion>,
mut scroll: EventReader<MouseWheel>,
) {
let Ok((mut cam, mut transform)) = query.get_single_mut() else { return };
let Ok((mut cam, mut transform, mut projection)) = query.get_single_mut() else { return };
let mut delta_motion = Vec2::ZERO;
for ev in mouse_motion.read() {
@@ -99,13 +192,20 @@ fn orbit_camera(
// Zoom (scroll)
cam.radius -= scroll_delta * cam.zoom_sensitivity * cam.radius;
cam.radius = cam.radius.clamp(0.5, 50.0);
cam.radius = cam.radius.clamp(cam.min_radius, cam.max_radius);
// Reset (R key)
if keys.just_pressed(KeyCode::KeyR) {
*cam = OrbitCamera::default();
}
// Keep the clip planes proportional to the zoom distance so depth precision
// stays usable across scales (tiny prop → whole stage grid).
if let Projection::Perspective(p) = projection.as_mut() {
p.near = (cam.radius * 0.02).clamp(0.02, 50.0);
p.far = (cam.radius * 50.0).max(2000.0);
}
*transform = orbit_transform(&cam);
}

File diff suppressed because it is too large Load Diff

View File

@@ -72,9 +72,28 @@ pub fn run() {
}
fn setup_scene(mut commands: Commands) {
commands.spawn(DirectionalLight {
illuminance: 10_000.0,
shadows_enabled: true,
..default()
// Bright ambient so surfaces facing away from every light aren't pure black.
commands.insert_resource(AmbientLight {
color: Color::srgb(0.9, 0.93, 1.0),
brightness: 900.0,
});
// A multi-directional rig (key + fills + rim + underside) so an orbiting
// camera always has some light on the visible side — the single front light
// left the backside unreadable.
let lights = [
(Vec3::new(1.0, 2.0, 1.5), 10_000.0, true), // key: front-top-right
(Vec3::new(-2.0, 1.0, 0.5), 4_500.0, false), // fill: left
(Vec3::new(0.5, 0.8, -2.0), 6_000.0, false), // rim: behind
(Vec3::new(0.0, -1.5, 0.5), 2_500.0, false), // underside fill
];
for (from, lux, shadows) in lights {
commands.spawn((
DirectionalLight {
illuminance: lux,
shadows_enabled: shadows,
..default()
},
Transform::from_translation(from).looking_at(Vec3::ZERO, Vec3::Y),
));
}
}

View File

@@ -10,10 +10,13 @@ use bevy::prelude::*;
use bevy_egui::{egui, EguiContexts};
use crate::iso_loader::{
FileInfo, FileSelected, ImageRgba, IsoState, PakContent, PakView, RequestOpenDir,
RequestOpenIso, SkyboxPreview, TextPreview, TexturePreview, VideoPreview, IsoLoaderSystemSet,
AudioPreview, FileInfo, FileSelected, ImageRgba, IsoState, ModelPreview, MovieSubtitles,
MovieVoice, PakContent, PakView, RequestAudio, RequestOpenDir, RequestOpenIso, RequestSubtitles,
RequestVoiceLibrary, SkyboxPreview, TextPreview, TexturePreview, VideoPreview, VoiceLibrary,
IsoLoaderSystemSet,
};
use crate::ViewerState;
use sylpheed_formats::SubLang;
pub struct ViewerUiPlugin;
@@ -123,6 +126,18 @@ fn render_dir(
}
}
/// Bundled event writers for the UI, to keep `draw_viewer_ui` under Bevy's
/// 16-parameter system limit.
#[derive(bevy::ecs::system::SystemParam)]
struct UiEvents<'w> {
open_iso: EventWriter<'w, RequestOpenIso>,
open_dir: EventWriter<'w, RequestOpenDir>,
file_selected: EventWriter<'w, FileSelected>,
subtitles: EventWriter<'w, RequestSubtitles>,
audio: EventWriter<'w, RequestAudio>,
voice_lib: EventWriter<'w, RequestVoiceLibrary>,
}
fn draw_viewer_ui(
mut contexts: EguiContexts,
mut viewer: ResMut<ViewerState>,
@@ -132,11 +147,14 @@ fn draw_viewer_ui(
text_preview: Res<TextPreview>,
mut pak_view: ResMut<PakView>,
skybox: Res<SkyboxPreview>,
model: Res<ModelPreview>,
mut video: ResMut<VideoPreview>,
file_info: Res<FileInfo>,
mut open_iso_events: EventWriter<RequestOpenIso>,
mut open_dir_events: EventWriter<RequestOpenDir>,
mut file_selected_events: EventWriter<FileSelected>,
mut subtitles: ResMut<MovieSubtitles>,
mut movie_voice: ResMut<MovieVoice>,
mut audio: ResMut<AudioPreview>,
mut voice_lib: ResMut<VoiceLibrary>,
mut events: UiEvents,
) {
let ctx = contexts.ctx_mut();
@@ -147,11 +165,11 @@ fn draw_viewer_ui(
#[cfg(not(target_arch = "wasm32"))]
{
if ui.button("Open ISO disc image…").clicked() {
open_iso_events.send_default();
events.open_iso.send_default();
ui.close_menu();
}
if ui.button("Open extracted folder…").clicked() {
open_dir_events.send_default();
events.open_dir.send_default();
ui.close_menu();
}
ui.separator();
@@ -166,6 +184,17 @@ fn draw_viewer_ui(
);
});
ui.menu_button("View", |ui| {
if ui.button("🎙 Voice Lines…").clicked() {
voice_lib.open = true;
if !voice_lib.loaded && !voice_lib.loading {
voice_lib.loading = true;
events.voice_lib.send_default();
}
ui.close_menu();
}
});
ui.menu_button("Help", |ui| {
if ui.button("Controls…").clicked() {
// TODO: controls popup
@@ -177,6 +206,101 @@ fn draw_viewer_ui(
});
});
// ── Standalone voice-line browser (floating window) ───────────────────
if voice_lib.open {
let mut open = true;
egui::Window::new("🎙 Voice Lines")
.default_width(360.0)
.default_height(480.0)
.open(&mut open)
.show(ctx, |ui| {
if voice_lib.loading {
ui.horizontal(|ui| {
ui.spinner();
ui.label("Reading sounds.tbl…");
});
ctx.request_repaint();
} else if !voice_lib.loaded {
ui.label("Open a game source first.");
} else {
ui.horizontal(|ui| {
ui.label("Filter:");
ui.text_edit_singleline(&mut voice_lib.filter);
});
let f = voice_lib.filter.to_lowercase();
// Group the thousands of entries as directory → speaker so the
// list is navigable (e.g. browse `Voice` by character to find a
// cutscene's radio line). `name` is `<lang>\<dir>\<file>.slb`.
use std::collections::BTreeMap;
let mut groups: BTreeMap<&str, BTreeMap<&str, Vec<&sylpheed_formats::slb::VoiceClip>>> =
BTreeMap::new();
for c in &voice_lib.clips {
if !f.is_empty() && !c.name.to_lowercase().contains(&f) {
continue;
}
let dir = c.name.rsplit('\\').nth(1).unwrap_or("?");
groups
.entry(dir)
.or_default()
.entry(c.speaker.as_str())
.or_default()
.push(c);
}
let shown: usize = groups.values().flat_map(|s| s.values()).map(Vec::len).sum();
ui.label(
egui::RichText::new(format!("{shown} / {} clips", voice_lib.clips.len()))
.weak()
.small(),
);
ui.separator();
let filtering = !f.is_empty();
egui::ScrollArea::vertical().show(ui, |ui| {
for (dir, speakers) in &groups {
let dtotal: usize = speakers.values().map(Vec::len).sum();
egui::CollapsingHeader::new(format!("📁 {dir} ({dtotal})"))
.id_salt(("vdir", *dir))
.default_open(filtering)
.show(ui, |ui| {
for (speaker, clips) in speakers {
egui::CollapsingHeader::new(format!(
"{speaker} ({})",
clips.len()
))
.id_salt(("vspk", *dir, *speaker))
.default_open(filtering || clips.len() <= 6)
.show(ui, |ui| {
for c in clips {
ui.horizontal(|ui| {
if ui
.button("")
.on_hover_text(&c.name)
.clicked()
{
audio.generation =
audio.generation.wrapping_add(1);
audio.loading = true;
audio.active = true;
audio.name = c.display.clone();
events.audio.send(RequestAudio {
clip: c.name.clone(),
display: c.display.clone(),
movie: None,
generation: audio.generation,
});
}
ui.label(&c.display);
});
}
});
}
});
}
});
}
});
voice_lib.open = open;
}
// ── Left panel: file browser ──────────────────────────────────────────
egui::SidePanel::left("file_browser")
.resizable(true)
@@ -225,7 +349,7 @@ fn draw_viewer_ui(
force_open,
selected,
loading,
&mut file_selected_events,
&mut events.file_selected,
&files,
);
}
@@ -251,11 +375,43 @@ fn draw_viewer_ui(
});
// ── Central area ──────────────────────────────────────────────────────
if browser.selected.is_some() {
if audio.active {
// Standalone audio player takes over the central panel.
egui::CentralPanel::default().show(ctx, |ui| {
ctx.request_repaint();
draw_audio_player(ui, &mut audio);
});
} else if browser.selected.is_some() && model.active && !browser.loading {
// 3D model: draw the central panel TRANSPARENT so the real Bevy scene
// (mesh + orbit camera) shows through, with just an info overlay on top.
egui::CentralPanel::default()
.frame(egui::Frame::none())
.show(ctx, |ui| {
egui::Frame::none()
.fill(egui::Color32::from_black_alpha(160))
.inner_margin(egui::Margin::same(8.0))
.rounding(egui::Rounding::same(4.0))
.show(ui, |ui| {
ui.heading(&model.name);
ui.label(&model.info);
ui.colored_label(
egui::Color32::from_gray(160),
"3D preview · LMB orbit · RMB pan · scroll zoom · R reset",
);
ui.colored_label(
egui::Color32::from_gray(140),
"position + normal + UV from the XBG7 vertex declaration",
);
});
});
} else if browser.selected.is_some() {
egui::CentralPanel::default().show(ctx, |ui| {
if browser.loading {
// A file was just clicked — show immediate feedback while the
// background read + decode runs, instead of the previous item.
// Keep repainting so the spinner animates even if the app is in a
// reactive (idle) winit mode while the worker thread churns.
ctx.request_repaint();
let name = browser
.selected
.and_then(|i| browser.files.get(i))
@@ -268,7 +424,43 @@ fn draw_viewer_ui(
});
});
} else if video.active {
draw_video_player(ui, &mut video);
let mut lang_changed = false;
let mut play_voice_solo = false;
draw_video_player(
ui,
&mut video,
&mut subtitles,
&mut movie_voice,
&mut lang_changed,
&mut play_voice_solo,
);
if play_voice_solo {
if let Some(movie) = movie_voice.movie.clone() {
video.playing = false; // pause the video; audio takes over
audio.generation = audio.generation.wrapping_add(1);
audio.loading = true;
audio.active = true; // show the panel immediately (spinner)
audio.name = format!("VOICE_{movie}");
events.audio.send(RequestAudio {
clip: String::new(),
display: format!("VOICE_{movie}"),
movie: Some((movie.clone(), movie_voice.lang)),
generation: audio.generation,
});
}
}
if lang_changed {
if let Some(movie) = subtitles.movie.clone() {
subtitles.generation = subtitles.generation.wrapping_add(1);
subtitles.cues.clear();
subtitles.loading = true;
events.subtitles.send(RequestSubtitles {
movie,
lang: subtitles.lang,
generation: subtitles.generation,
});
}
}
} else if !skybox.faces.is_empty() {
draw_skybox_grid(ui, &skybox);
} else if pak_view.loaded {
@@ -806,7 +998,14 @@ fn fmt_time(secs: f32) -> String {
/// keyboard shortcuts (space = play/pause, ←/→ = skip 10 s). Playback state
/// lives in `VideoPreview`; the decode/audio engine reacts to it in
/// `advance_video_playback`.
fn draw_video_player(ui: &mut egui::Ui, video: &mut VideoPreview) {
fn draw_video_player(
ui: &mut egui::Ui,
video: &mut VideoPreview,
subs: &mut MovieSubtitles,
voice: &mut MovieVoice,
lang_changed: &mut bool,
play_voice_solo: &mut bool,
) {
// ── Keyboard shortcuts (consumed so focused widgets don't also act). ──
let (mut toggle_play, mut skip) = (false, 0.0_f32);
ui.input_mut(|i| {
@@ -838,6 +1037,47 @@ fn draw_video_player(ui: &mut egui::Ui, video: &mut VideoPreview) {
ui.separator();
ui.colored_label(egui::Color32::YELLOW, "no audio");
}
// Subtitle + voice controls (right-aligned): language, CC, and Voice.
ui.with_layout(egui::Layout::right_to_left(egui::Align::Center), |ui| {
let before = subs.lang;
egui::ComboBox::from_id_source("subtitle_lang")
.selected_text(subs.lang.label())
.show_ui(ui, |ui| {
for lang in SubLang::ALL {
ui.selectable_value(&mut subs.lang, lang, lang.label());
}
});
if subs.lang != before {
*lang_changed = true;
}
ui.toggle_value(&mut subs.enabled, "CC")
.on_hover_text("Show subtitles");
if subs.loading {
ui.spinner();
} else if subs.enabled && subs.cues.is_empty() {
ui.weak("(no subtitles)");
}
ui.separator();
// Voice-over track: the movie's own stream carries only music+SFX, so
// this layers in the localized cutscene voice.
ui.toggle_value(&mut voice.enabled, "🗣 Voice")
.on_hover_text("Play the cutscene voice track over the movie");
if voice.loading {
ui.spinner();
} else if !voice.available {
ui.weak("(no voice)");
}
if voice.available
&& ui
.button("🎧")
.on_hover_text("Listen to the voice track on its own")
.clicked()
{
*play_voice_solo = true;
}
});
});
ui.separator();
@@ -849,6 +1089,7 @@ fn draw_video_player(ui: &mut egui::Ui, video: &mut VideoPreview) {
let scale = (avail.x / w).min(frame_h / h);
let (disp_w, disp_h) = (w * scale, h * scale);
let mut frame_rect = None;
if let Some(egui_id) = video.egui_id {
ui.allocate_ui(egui::vec2(avail.x, frame_h), |ui| {
ui.vertical_centered(|ui| {
@@ -860,10 +1101,26 @@ fn draw_video_player(ui: &mut egui::Ui, video: &mut VideoPreview) {
if resp.clicked() {
video.playing = !video.playing;
}
frame_rect = Some(resp.rect);
});
});
}
// Caption overlay: every cue active at the current position, stacked over
// the lower third of the frame (overlapping spans show together, newest —
// last in start order — lowest).
if let Some(rect) = frame_rect {
let active = subs.active_cues(video.position);
if !active.is_empty() {
let joined = active
.iter()
.map(|c| c.text.as_str())
.collect::<Vec<_>>()
.join("\n");
paint_caption(ui, rect, &joined);
}
}
// ── Transport bar ──
ui.horizontal(|ui| {
if ui.button(if video.playing { "" } else { "" }).clicked() {
@@ -907,6 +1164,86 @@ fn draw_video_player(ui: &mut egui::Ui, video: &mut VideoPreview) {
});
}
/// Draw a subtitle caption centered along the bottom of `rect`, with a
/// semi-transparent backing box so it stays legible over any frame.
fn paint_caption(ui: &egui::Ui, rect: egui::Rect, text: &str) {
let painter = ui.painter_at(rect);
// Font scales with the frame; clamped so it's readable but not huge.
let size = (rect.height() * 0.045).clamp(13.0, 30.0);
let font = egui::FontId::proportional(size);
let wrap = rect.width() * 0.9;
let galley = painter.layout(
text.to_string(),
font,
egui::Color32::WHITE,
wrap,
);
let margin = egui::vec2(10.0, 6.0);
let box_size = galley.size() + margin * 2.0;
let top_left = egui::pos2(
rect.center().x - box_size.x / 2.0,
rect.bottom() - box_size.y - rect.height() * 0.04,
);
let bg = egui::Rect::from_min_size(top_left, box_size);
painter.rect_filled(bg, 4.0, egui::Color32::from_black_alpha(160));
painter.galley(top_left + margin, galley, egui::Color32::WHITE);
}
/// Standalone audio player: a transport bar for a decoded voice/sound track with
/// no video. Playback state lives in `AudioPreview`; `advance_audio_playback`
/// reacts to it.
fn draw_audio_player(ui: &mut egui::Ui, audio: &mut AudioPreview) {
ui.horizontal(|ui| {
ui.heading(&audio.name);
ui.separator();
if audio.loading {
ui.spinner();
ui.label("decoding…");
} else {
ui.label("voice track");
}
ui.with_layout(egui::Layout::right_to_left(egui::Align::Center), |ui| {
if ui.button("✖ Close").clicked() {
audio.active = false;
}
});
});
ui.separator();
ui.add_space(ui.available_height() * 0.4);
// Big centered play/pause + a speaker glyph, since there's nothing to show.
ui.vertical_centered(|ui| {
ui.label(egui::RichText::new("🔊").size(64.0));
});
ui.add_space(12.0);
// Transport bar.
ui.horizontal(|ui| {
if ui
.button(if audio.playing { "" } else { "" })
.clicked()
{
audio.playing = !audio.playing;
}
ui.label(fmt_time(audio.position));
let tl_width = (ui.available_width() - 170.0).max(80.0);
ui.spacing_mut().slider_width = tl_width;
let mut pos = audio.position;
let resp = ui.add(
egui::Slider::new(&mut pos, 0.0..=audio.duration.max(0.1)).show_value(false),
);
if resp.changed() {
audio.position = pos;
audio.seek_request = Some(pos);
}
ui.label(fmt_time(audio.duration));
ui.separator();
ui.label("🔊");
ui.spacing_mut().slider_width = 90.0;
ui.add(egui::Slider::new(&mut audio.volume, 0.0..=1.0).show_value(false));
});
}
// ── World cubemap (skybox) viewer ─────────────────────────────────────────────
/// Show a world cubemap's 6 faces in a labelled grid (D3D9 order). A true

View File

@@ -0,0 +1,95 @@
# Handoff — movie subtitles, voice, and the movie manifest (2026-07-19)
Branch `feature/ipfb-idxd-parser`. This commit flushes several sessions of local
WIP; the **new, finished** work is the movie subtitle + voice + audio pipeline
and the movie manifest. One item is deliberately **left open** (see §4).
## 1. What's DONE and verified (static RE + tests)
### Movie subtitles (`crates/sylpheed-formats/src/movie_subtitle.rs`)
- Full movie→track→text chain (see `docs/re/structures/movie-subtitles.md`).
- **Multi-line caption fix**: a caption stored as several consecutive text tokens
sharing one timing (S13A: `"Look at it father"` + `"& beautiful isn't it"` @
`01:14.80`) — `parse_track` now joins them with `\n`. Previously all but the
last line were dropped. Disc test asserts both lines survive.
- **Overlap rendering**: `MovieSubtitles::active_cues(t)` returns every cue active
at `t`; the viewer stacks them (was: only the first cue shown).
- Umlauts / Latin-1 accents preserved (`utf16le_tokens`); Japanese label romanized
(CJK font still a known gap — see §5).
### Voice decode (`crates/sylpheed-formats/src/slb.rs`)
- `sound.pak` entries are XACT `.slb` banks wrapping **XMA1** (fmt tag `0x0165`,
48 kHz, mono content). `to_xma_riff` rebuilds a decodable RIFF; FFmpeg `xma1`
decodes it. Downmix mono via `-af pan=mono|c0=c0`.
- **Multi-subwave fix**: take the FIRST sub-wave bounded by its declared `data`
size — NOT `data..EOF` (which appended later takes = the S10S16 garble).
Verified: S13A→83.78s == video.
- `list_voice_clips` enumerates the voice library from `sounds.tbl`.
### Movie manifest (`crates/sylpheed-formats/src/movie_manifest.rs`) — the index
- `tables.pak` entry `ADVERTISE_MOVIE` (IDXD, hash `0x5B983A08`) is the
authoritative **movie → subtitle → voice** map. Located by shape (`is_manifest`),
parsed by grouping the string pool on `.wmv`.
- **Authoritative voice binding** replaces the old `VOICE_<movie>` guess:
101 movies, 83 with a `VOICE_` token, 18 without. The token is NOT always
`VOICE_<movie>` — 5 `hokyu_*` movies bind to `VOICE_D_450..454` in `\etc\`,
which a guess would miss. Token's subdir varies (Movie/etc/Voice) → resolve via
`sounds.tbl`. Disc test: all 83 resolved entries exist in `sound.pak`.
### Viewer wiring (`iso_loader.rs`, `ui.rs`)
- `resolve_movie_voice_clip` reads the manifest + `sounds.tbl` (DIRECT bindings
only) → `handle_voice_request` (movie player 🗣 toggle) and the 🎧 solo button
(`RequestAudio.movie`) use it. Unbound movies correctly post no audio.
- Standalone **View → 🎙 Voice Lines** browser (filter + ▶ play any `sound.pak`
clip) via `VoiceLibrary` + `AudioPreview`.
Tests: 46 formats-lib + 11 viewer-lib + disc (`movie_manifest_disc`,
`movie_subtitle_disc`, `slb_disc`) all green. Full `cargo test --workspace` green.
## 2. Also bundled in this commit (prior local WIP, not this session's focus)
- Viewer **async stage/model loading** (off-thread `prepare_xpr`, cancellable) —
`iso_loader.rs`/`ui.rs`/`camera.rs`. See memory `reborn-viewer-async-loading`.
- **Grouped-pool XBG7 / hero-ship** decode refinements in `mesh.rs`, `cli/main.rs`,
`mesh_disc.rs`. See memory `reborn-ship-drawlog` / `reborn-saber-texture-investigation`.
- `tools/analyze_drawlog_wvp.py`, `tools/extract_movie_subtitles.py`.
These build and test green but were not re-verified for behaviour this session.
## 3. USER TO VERIFY IN GUI (couldn't be tested from CLI)
1. S13A cutscene shows BOTH lines of the "Look at it father" caption.
2. Story-movie voice (S13A etc.) lines up with the video after the subwave fix.
3. The 5 directly-bound `hokyu_*` movies play voice; other hokyu are silent (see §4).
4. View → 🎙 Voice Lines browser lists and plays clips.
## 4. OPEN PROBLEM — unbound-hokyu resupply voice (do NOT re-guess)
The resupply cutscenes clearly SHARE voice recordings (video-only varies per
mission), but the correct join key for the **13 unbound `hokyu_*`** movies is
**unknown**:
- Manifest DIRECTLY binds only 5: `LS_s02A→450, LS_s09A→451, DS_s02A→452,
LS_s02H→453, DS_s07H→454` (grouping verified against the raw layout).
- Keying unbound movies by subtitle **demo id** (600→450…604→454, so
`hokyu_DS_s13A` demo602→`VOICE_D_452`) was implemented, tested against the
running game by the user, and is **WRONG** (wrong line). It has been **reverted**
— unbound hokyu now stay unvoiced rather than play wrong audio.
- Red flag: only `VOICE_D_450..454` exist; decoded durations `450=2.8s 451=1.6s
452=2.2s` but `453=0.14s 454=0.43s` — far too short for the spoken line, so
these `.slb` are likely **multi-subwave / not cleanly sliced** (same class as
the deferred B/C banks below).
**Next-box options:** (a) get the real join key from mission-event data (mission
scripts/IDXD tables that trigger resupply), or (b) build a proper multi-subwave
`.slb` decoder and verify audio by ear. Needs a concrete calibration clue from the
user (e.g. "s13A should sound like <movie X>" or the actual spoken words).
## 5. Other deferred items (unchanged)
- ~49 "layout-B/C" movie/voice banks (most RT + S07B/S11A/S12B/S13B) + some
individual clips need an XMA1 packet/seek-table reassembler.
- CJK/Japanese text rendering in the viewer (needs a CJK TTF via egui FontDefinitions).
- Texture colours (channel order / sRGB) still on the dynamic-RE backlog.
## 6. Resume checklist (next box)
1. `git pull` on `feature/ipfb-idxd-parser`; set `SYLPHEED_DISC=<extract root>`
(had `.../sylph_extract`).
2. Build guardrail: `CARGO_BUILD_JOBS=4` always (15 GB box; full `-j` OOM'd before).
3. `cargo test --workspace` (with `SYLPHEED_DISC`) should be green.
4. Pick up §4 (unbound-hokyu voice) — the only open thread from this session.

View File

@@ -0,0 +1,94 @@
# Handoff — XBG7 stage-container meshes, unified layout, mesh renderer (2026-07-12)
## Summary
This session extended the XBG7 mesh work from single weapons/props to the big
**stage containers**, fixed a long-standing weapon decode artifact, added a
headless mesh renderer for self-verification, and reworked the viewer's stage
presentation. Companion Xenia-Canary GPU instrumentation (the `log_draws`
draw-logger used as the ground-truth oracle) is on the `instrument-current`
branch of the Canary fork.
## What landed (Reborn — `feature/ipfb-idxd-parser`)
### 1. Stage containers decode — `Xbg7Model::stage_models`
`hidden/resource3d/Stage_S*.xpr` are **collections of up to ~450 enemy/prop
sub-models**, not single meshes. Each resource is a block
`[12-byte header][index buffer][vertex buffer]`; the index count is the
descriptor's `(index_bytes, index_count)` marker and the vertex count is the
`u32` 32 bytes before it. The blocks are **scattered among the container's
texture data with no stored offset**, so each is located by **content**: one
`O(file)` pass per stride finds vertex-buffer starts (unit NORMAL at vertex +12
whose previous stride slot isn't — a block boundary), then each resource is
pinned to the candidate where its indices validate and produce non-degenerate,
well-connected triangles. Fast (≤ ~4 s on a 70 MB stage), unambiguous.
- **Coverage: ~4993 sub-models across the 22 stages** (content-anchored).
- A **connectivity gate** (mean triangle edge ≤ 0.28× bbox diagonal) rejects
mis-anchored blocks that would render as spiky messes.
### 2. Weapon layout fix (unified with stages)
Weapons use the **same** `[12-byte header][index][vertex]` block layout. The old
decoder read the index buffer 12 bytes too early → 2 leading degenerate
triangles (the recurring "stray white triangle" artifact) + a lost tail of 6
real indices. Now the header is skipped; vertex offsets are unchanged, so
coverage stays 36/166 and the triangle lists are correct. Verified on
wep_00/03/04 via the renderer.
### 3. Headless mesh renderer — `sylpheed-cli mesh {info,render}`
A software rasterizer (orthographic, z-buffered, two-sided shading) that writes a
PNG. Lets anyone (including the assistant) verify recovered geometry without the
GUI. `--only <substr>` filters sub-models; `--yaw/--pitch/--size/--row` control
the view. Stage containers render as a normalised thumbnail grid.
### 4. Viewer UX
- Stage files show a **thumbnail grid**: each sub-model recentred + uniformly
scaled to a fixed cell (a 1-unit prop and a 3000-unit structure are equally
visible), skybox-plane / stray-anchor models (>20000 u) culled.
- Stage sub-models are **textured** by matching name → `_col` albedo.
- **Camera fixed**: zoom was hard-capped at 50 u (camera trapped inside big
stages); now zoom limits + clip planes scale to the framed model.
- Dispatch decodes both single-model and stage paths and keeps whichever yields
more geometry (fixes stages previously showing the dummy-root "black cube").
## State / how to verify
```bash
# formats + disc tests (needs the extracted disc)
SYLPHEED_RES3D=/path/to/hidden/resource3d \
cargo test -p sylpheed-formats --test mesh_disc -- --ignored --nocapture
# render any model / stage to a PNG
cargo run --release -p sylpheed-cli -- mesh render \
.../resource3d/Stage_S10.xpr /tmp/s10.png --size 1000
cargo run --release -p sylpheed-cli -- mesh render \
.../resource3d/rou_f001_wep_00.xpr /tmp/wep.png
# viewer
cargo run -p sylpheed-viewer # File → Open Directory → the extracted disc
```
## Known limitations / next steps
- **Multi-submesh main bodies with separate vertex pools** still mis-decode
(e.g. `Stage_S10` `e005`: a coherent fighter overlaid with wrong spike
triangles). `e003` works only because its two submeshes *share* one pool. Same
unsolved class as the **DeltaSaber hero body** (`f004`, still declined). Fix
needs per-submesh vertex-base RE. The LODs of these bodies (e.g. `e005_l`)
decode cleanly.
- **The per-resource data offset is not stored** — blocks are found by content.
Reversing the container's allocation order would give exact offsets + 100%
coverage but is a larger effort.
- **Texture colour correctness** remains the known "unverified" dynamic-RE item
(channel order / sRGB) for both 2D textures and model albedos.
- The viewer's stage decode runs on the main thread (~10 s on a 70 MB stage in a
debug build); move off-thread if the load hitch matters.
## Canary oracle (`instrument-current` on the Xenia-Canary fork)
`command_processor.cc::LogDrawForRE` (cvar `log_draws`) dumps each distinct
draw's primitive type, index buffer, vertex declaration, and sample vertex
positions to `xenia_re_draws.log`. Used to confirm the XBG7 layout (triangle
list, f32×3 position + f16×4 normal + f16×2 uv). Build with the memory-safe
flags (`CC=clang-19 CXX=clang++-19 CMAKE_BUILD_PARALLEL_LEVEL=3 … -j 3`); run one
emulator process at a time.

View File

@@ -14,11 +14,12 @@ Promote to a prose `structures/…md` file when a format needs behavioural notes
| name-hash (TOC keys) | ✅ | `sylpheed-formats/src/hash.rs` | Barrett-reduction hash; recovers original paths |
| IDXD object/table | ✅ | `sylpheed-formats/src/idxd.rs` | self-describing; ship/weapon stats verified vs known values |
| XPR2 texture + cubemap | 🟡 | `sylpheed-formats/src/texture.rs` | de-tile + A8R8G8B8; **colours unverified** (dynamic item) |
| T8aD 2D texture | 🟡 | `sylpheed-formats/src/t8ad.rs` | ~85% decode; **colours unverified**; ~15% variants deferred |
| T8aD 2D texture | 🟡 | `sylpheed-formats/src/t8ad.rs` | ~85% decode; **colours ✅ CONFIRMED** ([k8888](structures/texture-color-k8888.md)); ~15% variants deferred |
| RATC bundle | 🟡 | `sylpheed-formats/src/ratc.rs` | child listing confirmed; one level deep |
| LSTA sprite list | 🟡 | `sylpheed-formats/src/lsta.rs` | inline T8aD frames |
| IXUD subtitle | 🟡 | `sylpheed-formats/src/ixud.rs` | timed cues; **movie↔track link unknown** (dynamic item) |
| Fonts (ttf/otf/ttc) | ✅ | `sylpheed-formats/src/font.rs` | standard OpenType, parsed via ttf-parser |
| XBG7 mesh | 🟡/❔ | `sylpheed-formats/src/mesh.rs` + `tests/mesh_disc.rs` ([xbg7](structures/xbg7-mesh.md)) | weapons/props: declaration-driven variable stride (36 models), GPU-confirmed. **Stage containers: 5662 sub-models across 22 stages** via content-anchored grouped pools (`stage_models`). Quantized hero bodies (DeltaSaber `f004`) still declined |
## Functions / code paths

View File

@@ -0,0 +1,174 @@
# Movie subtitles & the movie ↔ mission ↔ text chain
Reverse-engineered 2026-07-19 (static, from the extracted disc). The full chain
that links a cutscene movie to its on-screen subtitle text is now closed.
## Files involved
- `dat/movie/*.wmv` — the cutscene videos. **Named by mission** (see below).
- `dat/movie/<lang>.pak` + `<lang>.p00` — per-language **subtitle timing tracks**
(`eng`, `jpn`, `deu`, `fra`, `esp`, `ita`) plus the caption font.
- `dat/GP_MAIN_GAME_<L>.pak` + `.p00` — per-language **caption TEXT**
(`E`=Eng, `J`=Jpn, `D`=Deu, `F`=Fra, `I`=Ita, `S`=Esp).
- `dat/tables.pak` entry `ADVERTISE_MOVIE` (hash `0x5B983A08`) — the master
**movie manifest**: all 101 `.wmv` names in mission-progression order, each
bound to its subtitle track and (optionally) its **voice bank** (see below).
## Movie filename → mission
Purely from the filename:
| Pattern | Meaning |
|---------|---------|
| `S<NN><P>.wmv` | Stage `NN` **story** cutscene, part `P` (A/B/C…) — e.g. `S02C` = stage 2, 3rd story scene |
| `RT<NN><P>.wmv` | Stage `NN` **radio / briefing** transmission, part `P` (`_1`/`_2` = split clips) |
| `hokyu_<LS\|DS>_s<NN><P>.wmv` | Stage `NN` **resupply** scene (`hokyu` = 補給). `LS`/`DS` = the two resupply-ship variants |
| `ADV.wmv` | Intro / title movie |
97 movies total: 27 story, 50 radio, 19 resupply, 1 intro. (Some referenced
stages — s24, s27, RT16 — exist as keys but the .wmv isn't in this extract.)
## Subtitle timing track — `<lang>.pak`
IPFB archive (`IPFB`, BE-u32 count, 16-byte header; TOC of
`[name_hash u32][offset u32][size u32]` triples, sorted by hash, into `.p00`).
- **Track key = `name_hash("subtitle_<movie_basename>.tbl")`** (the hash
lowercases internally, so basename case is irrelevant). This is the movie→track
link. Verified: `subtitle_S00A.tbl``0x6F2D9663`, `subtitle_hokyu_DS_s02A.tbl`
`0x3662B1F8`, `subtitle_RT01C_1.tbl``0x756F69FB`.
- The archive also holds **RATC** pre-rendered title-card / number textures
(`pwterop_s01a1.t32`, `pwrt_rt01_str.t32`, `pwnum0-9.t32`) + one TrueType font.
- **Each track data block is `Z1`+zlib**: bytes `5A 31` ("Z1"), a small header,
then a raw zlib stream (`78 DA`/`78 9C`). `zlib.decompress(blob[blob.find(b"\x78\xda"):])`.
- Decompressed = an **IXUD** container. Payload (UTF-16LE) is the timing sheet:
`SUBTITLE MSG_DEMO_<demo> <mm:ss.cc> MSG_DEMO_<demo> <mm:ss.cc> …`.
So the track says *which* demo-message shows *when*, not the text itself.
## Caption text — `GP_MAIN_GAME_<L>.pak`
Same IPFB+`.p00`. Among its ~1119 entries, **32 blocks are `Z1`+zlib → IXUD**
string containers holding the movie caption text. Layout: `IXUD`, u32 version,
hash@0x08, count@0x14, then `(recordhash,offset,len)` triples, then a UTF-16LE
string region where **each line is stored as `text` immediately followed by its
key** `MSG_DEMO_<demo>_<page>_<line>` (captions wrap across `_00`,`_01`, …).
537 English lines recovered. Entries 1 & 19 are the IDXD *schema* records
(`ID`, `PageCount`, `Character`=speaker e.g. `TCAFSUPPLY`, `Face`, line refs) —
no text, just structure.
## Movie → voice track: the manifest binding (`ADVERTISE_MOVIE`)
The `ADVERTISE_MOVIE` manifest is **also the authoritative movie→voice index**.
Its string pool emits, per movie, a run led by `<movie>.wmv` optionally followed
by `<pak>+….prt` (overlay art), `<pak>+SUBTITLE_<movie>.tbl`, and a bare
`VOICE_<token>`. Grouping the pool on `.wmv` (records are emitted in order)
recovers `movie → Option<voice_token>` without decoding the IDXD record binary
(`crate::movie_manifest`).
The voice token is **not** always `VOICE_<movie>`, so the manifest is required —
guessing both misses real bindings and invents tracks for silent movies:
- **83 / 101 movies have a voice token.** Story/radio movies use `VOICE_<movie>`
in `<lang>\Movie\`.
- **5 `hokyu_*` resupply movies bind to in-mission radio clips** — e.g.
`hokyu_LS_s02A → VOICE_D_450`, which lives in `<lang>\etc\`, *not* `Movie`.
A `VOICE_<movie>` guess would never find these.
- **18 movies have no direct voice token** = 4 boot logos + 1 HD test pattern +
**13 `hokyu_*` movies** (incl. `hokyu_DS_s13A`). Only the manifest's **direct**
bindings are trusted for playback.
### Shared resupply voice — UNRESOLVED for unbound movies
The manifest directly binds only **5** resupply movies, each to a shared
`VOICE_D_45x` clip in `<lang>\etc\`:
| bound movie | subtitle demo | clip |
|-------------|---------------|------|
| `hokyu_LS_s02A` | 600 | `VOICE_D_450` |
| `hokyu_LS_s09A` | 601 | `VOICE_D_451` |
| `hokyu_DS_s02A` | 602 | `VOICE_D_452` |
| `hokyu_LS_s02H` | 603 | `VOICE_D_453` |
| `hokyu_DS_s07H` | 604 | `VOICE_D_454` |
The resupply cutscenes clearly **share** voice recordings (only the video varies
per mission), so the 13 unbound `hokyu_*` movies must reuse one of these — but
the **correct join key is not yet known**:
- Keying by subtitle **demo id** (so `hokyu_DS_s13A`, demo 602 → `VOICE_D_452`)
was tried and is **WRONG** — it plays the wrong recording in-game. Do not use.
- Only `VOICE_D_450..454` exist (no 44x/45x neighbours). Decoded durations are
suspicious — `450`=2.8s, `451`=1.6s, `452`=2.2s, but `453`=**0.14s**,
`454`=**0.43s** — far too short for the spoken line, so these `.slb` banks are
likely **multi-subwave / not cleanly sliced** by the current extractor (same
class as the deferred B/C banks).
⇒ The unbound-hokyu voice mapping is **open** (needs either the real join key
from mission data, or a proper multi-subwave `.slb` decode + audio verification).
`hokyu_LS_s24A`/`s27A` have no subtitle track at all (stages absent from this
extract).
The token's `sound.pak` **subdirectory is not fixed** (`Movie` / `etc` / `Voice`),
so resolve it via `sounds.tbl` (which lists the full `<lang>\…\<token>.slb` path)
rather than assuming a directory. `movie_manifest::resolve_voice_entry` does the
full chain. Verified: all 83 resolved entries exist in `sound.pak`.
## Caption packing quirks (parser must handle)
- **Multi-line captions are split into consecutive text tokens** that share one
trailing timing, e.g. S13A stores `"Look at it father"` + `"& beautiful isn't
it"` before `01:14.80-01:17.60`. Accumulate every text token since the last
timing and join with `\n`; pairing strictly 1:1 silently drops all but the last
line.
- **Overlapping spans**: some tracks show two captions at once (an open-ended
radio line still up when the next range line starts). The viewer stacks every
cue active at `t` (`MovieSubtitles::active_cues`) instead of showing only the
first.
## The full join
```
tables.pak / ADVERTISE_MOVIE → list of movies (mission order)
<movie> ─ nh("subtitle_<movie>.tbl") ─→ <lang>.pak track
track (IXUD) → [ (MSG_DEMO_<d>, timecode), … ]
MSG_DEMO_<d> → GP_MAIN_GAME_<L> → "the localized caption line(s)"
```
## Coverage
- 92 / 97 movies have a subtitle track.
- **66 movies carry timed `MSG_DEMO` captions** — the **radio (`RT*`)** and
**resupply (`hokyu_*`)** movies. These fully decode to timed text.
- The **27 story (`S*`) movies have a track but 0 timed captions** — their text
is delivered as the **pre-rendered title-card textures** (`pwterop_*`, burned
styling), not MSG_DEMO lines.
## Worked examples (English)
```
RT01C_1.wmv (Stage 1 radio, part C):
00:00.50 [14] We did it! Okay, all pilots follow my lead!
00:06.80 [15] Rhino Leader to ACROPOLIS. We made it through and we're coming
home. Roger. It's good to see you're all safe.
00:19.30 [17] Yeah, but Brandon ... Damn. There's only seven of us. …
hokyu_DS_s02A.wmv (Stage 2 resupply):
00:00.00 [602] Resupply complete. You are cleared for take-off!
```
## Reusable extractor
`tools/extract_movie_subtitles.py` — parses `<lang>.pak`, resolves each movie's
track, cross-references `GP_MAIN_GAME_<L>` text, and prints per-movie timed
transcripts + the movie→mission table.
## In-mission dialogue (future work)
`GP_MAIN_GAME_<L>.pak` is the **global** message store, not just movie captions:
its `MSG_DEMO_*` table also holds the in-mission radio/dialogue lines (same demo
id space). So the *text* of gameplay dialogue is already decodable with
[`crate::movie_subtitle::build_demo_text`]. What's missing is the **trigger**
which demo id fires at which mission event — and that lives in the **mission
data** (mission scripts / `GP_MAIN_GAME` IDXD tables), not in the text pack. When
reversing mission data, look for demo-id references there to bind dialogue to
events; the IDXD "Message" schema records also carry `Character` (speaker) and
`Face` (portrait) per line.

View File

@@ -0,0 +1,62 @@
# Sound bank audio — `sound.pak` / `.slb` / XMA1
Reverse-engineered 2026-07-19 (static). `dat/sound.pak` (+ `sound.p00..p04`,
~1.08 GB) is the game's entire audio bank: **9519** IPFB entries, each an XACT
`.slb` sound bank wrapping **XMA1** audio.
## Names
Entry keys are `name_hash(<string>)` of filenames stored verbatim in the IDXD
string pools of `tables.pak``eng\sounds.tbl` (entry 39) and `jpn\sounds.tbl`
(entry 45). All 9519 recovered. Families:
- `<lang>\Voice\VOICE_<SPK>_<NNN>.slb` — per-line character voice (SPK = ADAN,
TCAF, RHIN, ADPL, BIRD, ACRO, ZZZZ…). Only `eng`/`jpn` voice exists.
- `<lang>\etc\VOICE_{A,B,C,D}_<NNN>.slb`, `<lang>\Briefing\BR<n>_<m>.slb`.
- **`<lang>\Movie\VOICE_<movie>.slb`** — the continuous voice track for cutscene
`<movie>.wmv` (e.g. `eng\Movie\VOICE_RT07A.slb`). Played from the video start.
- `BGM_###.slb` (32) — music. `Static.slb`, `JNGL_###` — SFX/ambience.
Only English and Japanese have voice; subtitles cover more languages, voice does
not. See [`crate::slb::movie_voice_name`].
## Codec
XMA1: RIFF `fmt ` tag **0x0165**, a 32-byte `XMAWAVEFORMAT`, **48000 Hz**. The
content is **mono** — some clips carry it in the **left channel only** (right is
silence), others are **dual-mono** (L = R). Downmix to mono by taking the left
channel (it always holds the full signal).
`XMAWAVEFORMAT` (one stream, 32-byte `fmt `): SampleRate @ fmt+16 (u32 LE),
Channels @ fmt+29 (u8), ChannelMask @ fmt+30 (u16). Movie voices report 2
channels / mask 0x2 despite the mono content.
## `.slb` layouts → decodable RIFF
Two on-disc layouts (`sound.pak` movie voices split ≈ 63 RIFF / 15 headerless):
- **RIFF present** — a `RIFF/WAVE` sits inside the bank; its 32-byte `fmt ` is
the real header. The XMA packets are **everything after that RIFF's `data`
chunk header** (`slb[data_pos+8..]`); the declared `data` size is unreliable
(sometimes half the real length), so take to end and clamp to the media length
downstream. Some banks have trailing bank data past the real audio — always
beyond the video length, so clamping to the video duration drops it.
- **Headerless** — no `RIFF` anywhere; a fixed **1392-byte** header precedes raw
XMA1 packets (48 kHz, 2 channels). Data = `slb[1392..]`.
[`crate::slb::to_xma_riff`] rebuilds a standalone XMA1 `RIFF/WAVE` for either
layout, ready for FFmpeg's `xma1` decoder.
## Decode (FFmpeg)
`ffmpeg -i rebuilt.riff -af "pan=mono|c0=c0" -t <media_len> out.wav` — the RIFF
is fed to the `xma1` decoder, downmixed to the left channel, clamped to length.
Verified: `VOICE_S00A` → 93.9 s == `S00A.wmv` 93.87 s; `VOICE_ADV` → 137.3 s ==
137.7 s; `VOICE_RT07A` → 49 s ≈ 50 s.
## Reading one entry cheaply
`sound.pak` is ~1 GB. Use [`crate::PakArchive::parse_toc`] on the small `.pak`
index, binary-search the name-hash, then read just `[offset, offset+size)` from
the `.pNN` segments (the viewer's `read_segment_range` seeks into the right
segment) instead of loading the whole archive.

View File

@@ -21,18 +21,30 @@ is sampled as **raw UNORM bytes — apply no sRGB/gamma decode.**
*Implication:* if the viewer/engine applies an sRGB→linear (or linear→sRGB) step to these, that is
wrong. Show the bytes as-is (raw RGBA8), gamma only where the game explicitly requests it.
### ❔ HYPOTHESIS — exact source **channel order** (needs measurement, do not infer)
Our decoders currently assume on-disk **A8R8G8B8** and reorder `[a,r,g,b] → [r,g,b,a]`. Canary
proves the *host* target is R8G8B8A8 identity-swizzled, but the *guest→host* byte order also depends
on the texture's runtime **endian** field (e.g. `8in32`), which is set by the game at fetch time and
**cannot be read statically** from the format table. So whether our reorder matches is unverified.
*What would confirm/refute it:* one **measured** decoded texel — either capture Canary's
framebuffer over a texture with distinct R≠G≠B pixels of known colour, or patch Canary's already-
present-but-unwired `TextureDump` (`src/xenia/gpu/texture_dump.cc`) to dump the decoded texture and
diff bytes against our decode. **Measure, don't guess** — a symmetric colour (white logo) can't
disambiguate channel order.
### ✅ CONFIRMED — on-disk order is **ARGB** (alpha first); `[a,r,g,b]→[r,g,b,a]` is correct
Settled by two independent lines that agree, without an emulator run:
1. **Measured on the disc data** — across 789 real T8aD textures (GP_HANGAR_ARSENAL, GP_MAIN_GAME_E,
DefTables), the byte position that is most-often exactly `0xFF` (the opaque-alpha signature of UI
art) is **byte 0 in 767/789 (97%)** — tf1 385/0, tf2 382/12. So texel byte 0 = alpha ⇒ on-disk
layout is **A,R,G,B**. (An earlier bimodality test tied byte0/byte3 only because colour channels
are also `0x00`-heavy from black backgrounds; the `==0xFF` discriminator isolates alpha cleanly.)
2. **Cross-check vs our Canary-validated renderer**`xenia-rs`'s `decode_k8888_tiled` (the path
behind the user-confirmed M1 splash / M2 video colours) nets memory-`ARGB` → endian-swap →
`swap(0,2)` = `[A,R,G,B]→[R,G,B,A]`, identical to ours. Since that renderer was confirmed against
Canary's output, Canary's ground truth is transitively in this chain.
Net: our channel order and the linear/UNORM interpretation are both **correct** for the ~85% RGBA
variants. A live Canary framebuffer capture would be redundant confirmation, not a new signal.
## Remaining / adjacent
- **XPR2** shares the ARGB→RGBA reorder + the linear/UNORM finding, but its **tiling** (de-tile) is a
separate path; if XPR2 skyboxes still look odd it's tiling or viewer display-gamma, not channel order.
- **Viewer display gamma:** decode is correct; if egui shows these too dark/bright it's how egui
interprets the RGBA8 (sRGB-encoded vs linear) at display time — a rendering nit, not a decode bug.
## Evidence log
- `2026-07-11` — static read of Canary `src/xenia/gpu/vulkan/vulkan_texture_cache.cc` host-format
table + confirmed no `R8G8B8A8_SRGB` entry exists. → sRGB/UNORM claim `CONFIRMED`; channel-order
claim remains `HYPOTHESIS` pending a measured texel.
- `2026-07-11` — static read of Canary `vulkan_texture_cache.cc` host-format table; no `R8G8B8A8_SRGB`
entry exists → sRGB/UNORM claim `CONFIRMED`.
- `2026-07-11` — measured `==0xFF` alpha dominance over 789 disc T8aD textures (byte0 = alpha 97%) +
cross-checked vs `xenia-rs decode_k8888_tiled` (the Canary-validated M1/M2 path). → channel-order
claim promoted `HYPOTHESIS → CONFIRMED`.

View File

@@ -0,0 +1,254 @@
# XBG7 — mesh geometry (inside XPR2 model containers)
- **Confidence:** 🟡 `PROBABLE` for the single-stream layout (below); ❔ `HYPOTHESIS` / undecoded
for the complex multi-stream body layout.
- **Parser in:** `sylpheed-formats/src/mesh.rs` (`Xbg7Model::from_xpr2`), tests
`tests/mesh_disc.rs`. Container parsing reused from `src/texture.rs` (`Xpr2Header` /
`Xpr2ResourceEntry`).
- **Applies to:** ship / weapon / prop models in `hidden/resource3d/*.xpr` (166 files).
- **Method:** clean-room — static hex inspection of the retail disc + geometric validation of the
recovered triangles (non-degenerate area, indices in range, bbox matches the descriptor's stored
size). No game code decompiled or copied.
## Where XBG7 lives
Models are ordinary **`XPR2`** containers (see the XPR2 texture doc / `texture.rs`). The 16-byte
resource-directory entries (from file offset `0x10`) carry `TX2D` texture resources **and** one or
more `XBG7` geometry resources:
```
entry = [ tag:4 ][ data_offset:u32 ][ descriptor_size:u32 ][ name_offset:u32 ] (big-endian)
```
Offsets are relative to the directory base `0x10`. The `XBG7` *descriptor* (at `data_offset+0x10`,
`descriptor_size` bytes) is a **scene / material / node graph** — it holds node names
(`rou_f001_mnt1_root`, `Light`, …), a bounding value (`0x41F00000` = 30.0 ≈ the ship's ~30-unit
length), material names matching the `TX2D` channels (`_col` albedo, `_spc` specular, `_gls` gloss,
`_lum` luminance), and per-sub-mesh records. The **vertex / index buffers** live in the container's
shared **data section** (from `header_size`).
## Sub-mesh records (in the descriptor)
Read in file order by a sliding 4-byte scan; each is a big-endian tuple:
```
[ vtx_count:u32 ][ 0:u32 ][ idx_count:u32 ][ tail:u32 ]
3..=65535 ==0 mult. of 3 1..=64
```
(For `rou_f001_wep_00`: `vtx_count=215`, `idx_count=1092` — matches the recovered geometry exactly.)
## The single-stream data layout — 🟡 PROBABLE (decoded, GPU-cross-checked)
For 36 of the 166 models (weapons, simple props) the data section is a straight sequence of
sub-meshes, carved from `header_size` in record order:
```
per sub-mesh block:
[ 12-byte header (contents undecoded) ]
[ index buffer : idx_count × u16 BE ] triangle list (prim=4, GPU-confirmed)
[ vertex buffer : vtx_count × stride bytes ] ← declaration-driven
(pad to 16 bytes → next sub-mesh)
```
The 12-byte header precedes the **index** buffer (the same block shape as stage
resources — see below); the vertex buffer follows the indices with no further
gap. (Earlier this was mis-modelled as `[index][12-byte gap][vertex]`, which put
the vertex buffer at the identical offset but read the index buffer 12 bytes too
early — turning the 12 header bytes into 6 junk indices = **2 leading degenerate
triangles** (a stray-triangle artifact) and dropping the last 6 real indices.
Skipping the header fixes the triangle list with no change to vertex coverage.)
**Vertex declaration.** The layout is **not fixed-stride**. The descriptor holds a declaration table
(right after the `(index_bytes, index_count)` marker) of `{offset:u32, format-code:u32, usage<<16:u32}`
big-endian triples, terminated by `offset == 0x00FF0000` / `code == 0xFFFFFFFF`:
| usage | element | format code | format | size |
|-------|----------|-------------|--------|------|
| 0x00 | POSITION | `0x2A23B9` | f32×3 | 12 B |
| 0x03 | NORMAL | `0x1A2360` | f16×4 (use xyz) | 8 B |
| 0x05 | TEXCOORD | `0x2C235F` | f16×2 (u,v) | 4 B |
Stride = max element extent. Models **omit elements** → variable stride (20 = pos+normal, no UV;
24 = pos+normal+uv). Each element is read in **naive big-endian component order** (the raw file
bytes; see the endianness note below). Assuming a fixed stride-24 was why the old decoder mis-aligned
and declined the pos+normal-only models.
**Alignment pinned by the normals.** The vertex buffer starts at `index_end + 12` (a fixed 12-byte
header — NOT `align16`, which lands 4 bytes early on most models). The correct offset is the unique
one where recovered normals are exactly unit-length.
**Safety gate:** every index is validated `< vtx_count`, and when the declaration has a normal
element the mean recovered `|normal|` must be ≈1 (`[0.5, 2.0]`). A model failing either is rejected
(`MeshError::UnsupportedLayout`) rather than emitting garbage.
### Endianness — file bytes are naive-BE, `k8in32` is a red herring ✅
The Canary GPU capture reports every vertex stream with fetch `endian = 2` (`k8in32`, each 32-bit
word byte-reversed). This describes the **guest-memory** copy the GPU fetches — **not** the `.xpr`
file bytes. Reading the file with a `k8in32` transform breaks the normals (mean `|normal|` → 1.33);
the plain per-element big-endian read yields exactly unit normals. So the game **rearranges** the
vertex data between the on-disc `.xpr` and the uploaded buffer; the decoder reads the file directly
and must use naive BE.
## Stage containers — multi-resource, grouped pools ✅ (content-anchored)
`hidden/resource3d/Stage_S*.xpr` are not single models but **collections of
enemy / prop sub-models** — up to ~400 `XBG7` resources each (e.g. `Stage_S07` =
378). Their data layout differs from the weapon files:
- Each resource is a block `[12-byte header][index buffer][vertex buffer]`.
Unlike weapons there is **no 12-byte gap between index and vertex** — the
vertex buffer directly follows the indices.
- The **index count** is the descriptor's `(index_bytes, index_count)` marker
(the *total* for the resource — a resource may have several sub-meshes summing
to it), **not** the first sub-mesh record. The **vertex count** is a `u32`
stored **32 bytes before** the marker.
- The blocks are **scattered through the data section, interleaved with the
container's texture data**, in an allocation order that is **not** directory
order and is **not stored** in any descriptor field we could find (the
descriptor holds sizes — `rel 160 ≈ index_bytes+3`, `rel 164 =
0x1000_0000 | (vertex_bytes+2)` — but no data offset). Reconstructing that
allocation order is unsolved.
Because the offset is not stored, each resource's block is located by
**content**: parser `Xbg7Model::stage_models` does one `O(file)` pass per
distinct stride to find every **vertex-buffer start** (an offset whose NORMAL —
`f16×4` at vertex `+12` — is unit length while the *previous* stride slot's is
not, i.e. a run boundary; ~one candidate per block, not millions), then pins each
resource to the unique candidate where the `index_count` indices ending just
before it are all `< vertex_count`, reference nearly all vertices, and yield
non-degenerate triangles with real extent. This is fast (≤ ~2 s on a 70 MB
stage) and unambiguous (no two resources collide). Blocks that fail validation —
the few quantized hero bodies — are **skipped**, never emitted as garbage.
**Coverage:** 5662 sub-models decode across the 22 stage files (e.g. `Stage_S07`
366/378, `Stage_S10` 7/9 — including the main enemy bodies `e003`/`e005`, their
LODs, weapons, and props). The viewer (`spawn_stage_models`) lays the decoded
sub-models out as a side-by-side "cast sheet", skipping the few huge skybox-plane
resources (300 k-unit quads). Cross-checked: `e003` = 2383 v / 1436 t bbox
23.5×9.5×32.5; `e005` = 2566 v / 1507 t; both 0 degenerate.
## Not yet decoded — ❔ the complex body layout
The hero-ship *body* meshes (`DeltaSaber_A/_T/_W.xpr` resource `f004`, and ~100 other models) still
decline. Two open sub-problems: (1) **multi-sub-mesh models** whose *first* sub-mesh decodes but a
later one's inter-mesh offset isn't yet handled (the `align16` advance is a guess) — these are
declined whole; (2) the big body meshes, where the data section does not start with an index buffer
and the geometry sits at descriptor-addressed offsets. NB the GPU capture showed **every** rendered
mesh is *single-stream* (just wider strides, e.g. 44 bytes = pos + f32×3 + colour + 2×f16×4), so the
body is likely single-stream-with-a-richer-declaration rather than the "separate streams" first
guessed — it was simply not rendered in the captured session (menu only). A capture taken *in a
mission* (where `DeltaSaber` renders) would hand over its exact declaration directly. `DeltaSaber_A`'s
5 XBG7 blocks are `f004` (body) + `_rou_f004_mnv01_L/_R`, `_mnv02`, `_turn180` (maneuver / pose).
## Evidence log
- 2026-07-12 — `rou_f001_wep_00.xpr`: XPR2 dir = 1×XBG7 (`f001_wep_00`) + 3×TX2D
(`_col/_gls/_spc`). Data section starts with a u16-BE index run (max 214), then stride-24
vertices. Descriptor record `[215,0,1092,4]` at desc `+0x2B0`; index-buffer byte size `0x888`
(=2184=1092×2) at desc `+0x1A0`. Recovered mesh = 215 v / 364 t, 362 non-degenerate, median tri
area 0.015 — coherent. → single-stream layout **PROBABLE**.
- 2026-07-12 — descriptor-driven sequential carving over all 166 models: **42 carve** under the
index-only check. Real 3D extent confirmed on `rou_f001_wep_04` (bbox 3.7×1.2×3.5).
- 2026-07-12 (refine) — attribute ranges differed per model (`wep_00` UV≈attr0/2, `wep_04`
attr4/5, `wep_03` attr1≈±60000 = garbage) → **refuted the fixed "6 half attrs, UV=attr0/2"
reading.** Found the descriptor **vertex declaration** (offsets 0x0C/0x14, usages 0x03 NORMAL /
0x05 TEXCOORD). Solved the vertex-base offset with a **unit-normal validator**: `index_end + 12`
gives median `|normal| = 1.000` on every weapon model (`wep_00/02/03/04`), vs `align16` landing 4
bytes early. UVs then land in `[0,1]`. Adding the normal gate: **25 models decode clean +
normal-valid** (the rest — incl. `Stage_S*` degenerate blobs — correctly declined).
- 2026-07-12 (DYNAMIC) — added a cvar-gated draw logger to Canary
(`command_processor.cc::LogDrawForRE`, cvar `log_draws`), captured the Ready Room / Briefings.
**GPU ground truth confirmed the static layout exactly**: primitive `prim=4` = **triangle LIST**
(settles the list-vs-strip question), and a stream with `f32x3 @offset0` + `f16x4 @3dw` +
`f16x2 @5dw`, stride 6 dwords = 24 bytes — matching `POSITION@0, NORMAL@0x0C, TEXCOORD@0x14`.
Revealed **stride varies** (24, 20, 28, 44 …) and that all streams are **single-stream**
parsed the declaration for variable stride: coverage **25 → 36** (e.g. `wep_05` is pos+normal,
stride 20, no UV — previously mis-aligned). Also confirmed the endianness note above: fetch
`endian=2` (k8in32) is the *guest* copy; file stays naive-BE. `Stage_S*` now decode (stride 20).
- 2026-07-12 (STAGE) — `Stage_S*.xpr` decoded as multi-resource containers. Found each geometry
block is `[12B hdr][index buffer][vertex buffer]`, index count = descriptor marker (total, e.g.
`e003` = 4308 spanning 2 sub-meshes), vertex count = `u32` at marker32 (`e003` = 2383 → verified
by max-index 2382 and 0 degenerate tris, bbox 23.5×9.5×32.5). Blocks are scattered among texture
data with **no stored offset** (descriptor rel 160 = idx_bytes+3, rel 164 = `0x1000_0000 |
(vtx_bytes+2)` are sizes, not offsets; block starts e.g. `e003`@0x10230, `e003_l`@0x5d000,
`e005_l`@0x9b000 are not directory-ordered). Solved by **content anchoring**: a single per-stride
pass finds vertex-run starts (unit NORMAL at +12 whose previous slot isn't), then match each
resource by strict index+triangle validation. **5662 sub-models decode across 22 stages** (S07
366/378), incl. the previously-declined main bodies `e005` (2566 v) and weapons — ≤2 s on 70 MB.
Parser `Xbg7Model::stage_models`, tests `stage_models_{decode,sweep,quality_audit}`.
- 2026-07-17 — **triangle-LIST re-confirmed; a strip interlude refuted; winding-consistency gate
added.** A 2026-07 change had briefly re-read the index buffers as triangle *strips* (to "fill
holes"). Refuted objectively with a new `XVERIFY` diagnostic that compares both readings by
**stored-normal agreement** (each triangle's cross-product face normal vs the sum of its vertices'
stored normals): the LIST reading gives agreement **1.000** on every clean weapon (`wep_00/03/04/19`
= only possible with correct topology + winding), the STRIP reading **~0.49** (random). The strip
reading also over-generated ~2.5× the triangles (wep_00: 938 vs 364) — a hole-filling garbage soup.
Reverted to LIST in both paths (`from_xpr2`, `read_pool_mesh`), matching the `prim=4` GPU capture.
Added an objective **winding-consistency gate** `max(na, 1-na)`: a correct carve is internally
consistent (agreement ≈1.0, or ≈0.0 for inverted-but-consistent winding — a real single-sided
mesh), a mis-carve scatters to the ≈0.5 middle. `from_xpr2` declines sub-meshes below 0.90 (e.g.
`wep_23` na=0.398 → declined instead of a spike-mess); the single-model **content-anchor fallback**
gates at 0.85; the large multi-resource **stage** path stays ungated (its enemy meshes span a
continuous 0.51.0 consistency range — a hard gate there dropped ~48/314 legit S07 blocks).
**Routing fixed:** `decode_models` (CLI) and the viewer now route by `count_xbg7` (1 → validated
records-based list decode, fallback to strict-gated anchor; >1 → stage anchor) instead of the old
"whichever decoder yields more verts" rule — that rule let stage content-anchoring win on
single-model weapon files and fabricate **phantom** blocks (a `wep_00` clone appearing inside
`wep_19`), duplicates, and spike-mess anchors. Weapons now: 33 clean-decode / 26 declined (declined
= genuinely multi-stream or un-carvable, shown as nothing rather than garbage); stage coverage
unchanged (S07 314). `expand_triangle_strip` retained as an `XVERIFY`-only diagnostic.
- 2026-07-12 — `DeltaSaber_A.xpr` body: data does **not** begin with indices; plain-`f32×3` runs
with ship-scale extent (span ≈2734, matching bbox 30.0) found only at high offsets
(`data+0x28634C`, …) → multi-stream, **undecoded**.
- 2026-07-18 — **GROUPED-POOL layout cracked → the hero ship (Delta Saber) fully decodes.** The
detailed models (`DeltaSaber_*.xpr` + ~100 others) were declined for **location**, not format —
their vertex format is the standard stride-24 triangle list. A resource's *several* sub-meshes
don't interleave `[idx][vtx]` per block; they share **two grouped pools**: an **index pool**
(buffers concatenated in descriptor-marker order, each **4-byte aligned**) followed by a **vertex
pool** (each sub-pool `vtx_count × stride`, same order), with the index pool ending **exactly**
where the vertex pool begins. So the whole resource pivots on one unknown, the first vertex-pool
start `vb0` (= index-pool end, found by the unit-normal vertex-run scan); everything else is
derived: `ib0 = vb0 span`, `ib[i] = align4(ib[i-1] + idx_count[i-1]·2)`,
`vb[i] = vb[i-1] + vtx_count[i-1]·stride`. Reversed statically from `DeltaSaber_T.xpr` and
**cross-checked against a Canary GPU draw-log capture** (mission ship = `DeltaSaber_T.xpr`, found
via the `--log_file_io` kernel hook): `f001` = body (idx@`data+0xC` = 0x5500C, vtx@0x61ACC, 10891 v
/ 8187 t) + **7 detail parts** (fins/cockpit/wingtips, markers at descriptor 0x3BEC…0x58FC) =
**8650 tris**, and **every sub-mesh decodes at 0 degenerate / full coverage / winding-agreement
1.000**. This is the layout the per-block adjacency anchor (`ib = vb idx_bytes`) rendered as a
**spiky phantom** (it read 24561 indices starting 2782 B too late, agree 0.64, 1277 degenerate).
Insight: a single index marker reduces the grouped model to `index_end = vb0`, i.e. the existing
adjacency `ib = vb idx_bytes` — so grouped **generalises** the single-block anchor (n=1 is
identical). Implemented as `anchor_grouped_meshes` (mesh.rs): `anchor_models` routes resources with
>1 index marker to it (validated per sub-mesh; on failure falls back to the old first-marker
adjacency anchor so stage coverage never regresses); single-marker stages/props keep the exact
prior path. The shared acceptance test is factored into `validate_block` (the connectivity
heuristic is relaxed for *derived* grouped parts, which are pinned by in-range + consistency, so
small flat fins aren't mis-rejected). Render self-check: `sylpheed-cli mesh render DeltaSaber_T.xpr
--only f001` (exact-name match excludes the `_rou_f001_mnv*` animation poses) → clean complete
fighter. Test `hero_ship_grouped_pool_decodes`. Colours/UVs still pending the running-game oracle.
- 2026-07-18 (refinement) — **4-byte vertex-pool alignment + weapon recovery.** The grouped-pool
rule "index pool ends exactly where the vertex pool begins" is really "the vertex pool is **4-byte
aligned** after the index pool": `vb0 = align4(ib0 + span)`, so 0..=3 bytes of padding can sit
between them. DeltaSaber's index pool ended already-aligned (pad 0), which hid this; **19
weapon/`*_hangar` models** (single- and multi-marker: `wep_08/11/34/58/62/69/81/83…`) have pad 2
and so decoded to *nothing* — the viewer then showed them as a flat 2D texture instead of a model.
Fix: both anchors try `pad ∈ 0..=3` (`ib = vb idx_bytes pad` for the single-block adjacency
anchor; `ib0 = vb0 span pad` for the grouped pivot), validated — a wrong pad reads shifted
indices → agreement collapses < 0.85, so only the true pad passes. pad>0 in the ungated stage path
is gated at a strict 0.85 to avoid a false anchor; pad 0 keeps its exact prior behaviour (stages
unchanged). Result: all 19 now decode as clean models (e.g. `wep_34` 1243 v / 1233 t, a
long-barrelled gun-pod; `wep_08` 3 sub-meshes / 478 t). Viewer routing already falls through
`from_xpr2``anchor_models(0.85)` for single-XBG7 files, so the recovered grouped/padded weapons
now preview as meshes.
- 2026-07-18 (refinement 2) — **pivot on the largest sub-mesh; all 19 recovered.** Three weapons
(`wep_81`, `wep_81_hangar`, `wep_30_hangar`) still declined because the grouped pivot validated
`markers[0]`, which for these is a tiny *elongated* lead bracket that fails the connectivity gate
even when perfectly placed. Fixed by pivoting the alignment check on the **largest** marker (max
index count) — the sub-mesh whose triangle-quality/connectivity signature most reliably confirms
`(ib0, vb0)`. Once the pivot validates, markers up to it are read unconditionally (a legitimately
tiny/flat lead part may fail the quality gates yet still be real), and markers after it stay
validated so a stray trailing marker ends the chain. Result: **all 19 previously-declined weapons
decode** (`wep_81` 460 t missile w/ tail fins; `wep_30_hangar` 334 t). DeltaSaber unchanged (its
body IS the largest marker → same pivot). 7/7 disc tests green, stage quality audit unchanged.

View File

@@ -0,0 +1,126 @@
#!/usr/bin/env python3
"""Recover each ship part's runtime WORLD transform from a Canary draw log.
The draw logger (branch capture-drawlog) now dumps, per distinct draw:
- the first ~8 object-space vertex positions (pre-transform), and
- the vertex-shader float constants c0..c31 (which hold the transform matrices).
The game bakes each part's WORLD matrix into the constants, so the shader's
World*View*Proj (WVP) for a part is VP * part_World. The body (bdy_01) is drawn
with an identity world, so body_WVP = VP. Hence for any part drawn in the SAME
frame (the ship's parts always are):
part_World = inv(body_WVP) * part_WVP (View/Proj cancel)
We don't know a-priori which 4 consecutive constants form the WVP, so we try every
block c[i..i+3] and keep the one for which inv(body_WVP)*part_WVP is a RIGID
transform (orthonormal 3x3 + translation) for the parts — that uniquely identifies
the matrix block. The translation column of each part's world matrix is the
placement we want (esp. the ±X nacelle offset the static XBG7 graph lacks).
Usage: analyze_drawlog_wvp.py xenia_re_draws.log
Identify the body draw via --body-indices (the draw with the most indices) or by
matching object-space positions to our decoded sub-meshes.
"""
import sys, re
import numpy as np
def parse(path):
draws = []
cur = None
consts = {}
mode = None
for line in open(path):
if line.startswith("DRAW "):
if cur is not None:
cur["consts"] = consts
draws.append(cur)
cur = {"header": line.strip(), "positions": [], "textures": [],
"size_words": 0}
consts = {}
mode = None
m = re.search(r"index\[base=0x([0-9A-Fa-f]+) count=(\d+)", line)
cur["idx_base"] = int(m.group(1), 16) if m else None
cur["idx_count"] = int(m.group(2)) if m else None
m = re.search(r"indices=(\d+)", line)
cur["indices"] = int(m.group(1)) if m else None
elif cur is not None and line.strip().startswith("stream "):
# Largest vertex buffer (size_words) = the hull (body, identity world).
m = re.search(r"size_words=(\d+)", line)
if m:
cur["size_words"] = max(cur["size_words"], int(m.group(1)))
elif cur is None:
continue
elif line.strip().startswith("positions:"):
cur["positions"] = [tuple(map(float, t.split(",")))
for t in re.findall(r"\(([^)]+)\)", line)]
elif line.strip().startswith("textures:"):
cur["textures"] = re.findall(r"base=0x([0-9A-Fa-f]+)", line)
elif line.strip().startswith("vsconst"):
mode = "vsconst"
elif mode == "vsconst":
m = re.match(r"\s*c(\d+)\s+(-?[\d.]+)\s+(-?[\d.]+)\s+(-?[\d.]+)\s+(-?[\d.]+)", line)
if m:
consts[int(m.group(1))] = np.array([float(m.group(i)) for i in range(2, 6)])
else:
mode = None
if cur is not None:
cur["consts"] = consts
draws.append(cur)
return draws
def wvp_from(consts, i):
"""4x4 from constants c[i..i+3] as rows."""
if not all(k in consts for k in (i, i+1, i+2, i+3)):
return None
return np.array([consts[i], consts[i+1], consts[i+2], consts[i+3]])
def is_rigid(M):
"""3x3 upper-left orthonormal (rotation/reflection), within tolerance."""
R = M[:3, :3]
if not np.all(np.isfinite(R)):
return False
G = R @ R.T
return np.allclose(G, np.eye(3), atol=1e-2) or np.allclose(G / (np.trace(G)/3), np.eye(3), atol=1e-2)
def main():
if len(sys.argv) < 2:
print(__doc__); return
draws = parse(sys.argv[1])
print(f"parsed {len(draws)} draws")
# Body = the draw with the largest vertex buffer (the ~10891-vert hull, drawn
# with an identity world). idx_count is unreliable (the hull's 9 material
# groups dedup to the first group's partial count).
body = max(draws, key=lambda d: d.get("size_words") or 0)
print(f"body draw: {body['header'][:90]} size_words={body['size_words']}")
# Try every constant block; report the one giving rigid part transforms.
for i in range(0, 29):
bwvp = wvp_from(body["consts"], i)
if bwvp is None:
continue
try:
binv = np.linalg.inv(bwvp)
except np.linalg.LinAlgError:
continue
worlds = []
rigid = 0
for d in draws:
pw = wvp_from(d["consts"], i)
if pw is None:
continue
W = binv @ pw # column-vector convention; may need transpose
worlds.append((d, W))
if is_rigid(W) or is_rigid(W.T):
rigid += 1
if rigid >= max(2, len(worlds)//2):
print(f"\n=== WVP block c{i}..c{i+3}: {rigid}/{len(worlds)} rigid ===")
for d, W in worlds:
# translation is the last column (or last row, convention-dependent)
tcol = W[:3, 3]
trow = W[3, :3]
print(f" idx={d['idx_count']:>6} base={d['idx_base'] and hex(d['idx_base'])}"
f" t_col=({tcol[0]:+.2f},{tcol[1]:+.2f},{tcol[2]:+.2f})"
f" t_row=({trow[0]:+.2f},{trow[1]:+.2f},{trow[2]:+.2f})")
if __name__ == "__main__":
main()

View File

@@ -0,0 +1,88 @@
#!/usr/bin/env python3
import struct, zlib, re, os
MOD=0x00FFF9D7; RECIP=0x80031493
def rol(v,n,b=32):
v&=(1<<b)-1; return ((v<<n)|(v>>(b-n)))&((1<<b)-1)
def nh(name):
bs=bytearray(name.encode('latin-1'))
for i,b in enumerate(bs):
if 0x41<=b<=0x5A: bs[i]=b+0x20
a=0;bb=0
for byte in bs:
c=(byte if byte<0x80 else byte-0x100)&0xFFFFFFFF
a=((rol(a,8)&0xFFFFFF00)+c)&0xFFFFFFFF
bb=(bb+c)&0xFFFFFFFF
hi=((a*RECIP)>>32)&0xFFFFFFFF; q=rol(hi,9)&0x1FF
a=(a-(q*MOD))&0xFFFFFFFF
return (((bb<<24)&0xFF000000)|(a&0xFFFFFF))&0xFFFFFFFF
EX="/home/fabi/RE Project Sylpheed/sylph_extract"
def load_pak(base):
"""base like dat/movie/eng ; returns list of (hash,off,size), data bytes"""
pak=open(f"{EX}/{base}.pak","rb").read()
n=struct.unpack(">I",pak[4:8])[0]
es=[struct.unpack(">III",pak[16+12*i:28+12*i]) for i in range(n)]
p00=f"{EX}/{base}.p00"
data=open(p00,"rb").read() if os.path.exists(p00) else pak
return es, data
def unz(blob):
if blob[:2]==b'Z1':
zi=blob.find(b'\x78\xda');
if zi<0: zi=blob.find(b'\x78\x9c')
return zlib.decompress(blob[zi:])
return blob
# --- verify the track key scheme ---
es,data=load_pak("dat/movie/eng")
keys={h for h,_,_ in es}
for mv,exp in [("S00A",0x6F2D9663),("hokyu_DS_s02A",0x3662B1F8),("RT01C_1",0x756F69FB),("S16A",0xEEB85377)]:
k=nh(f"subtitle_{mv}.tbl")
print(f"subtitle_{mv}.tbl -> {k:08X} expect {exp:08X} inpak={k in keys} {'OK' if k==exp else 'MISMATCH'}")
# --- build MSG_DEMO text table from GP_MAIN_GAME_E ---
es2,data2=load_pak("dat/GP_MAIN_GAME_E")
text={} # 'MSG_DEMO_ddd_ppp_ll' -> string
keyrx=re.compile(r'^MSG_DEMO_\d+(_\d+)*$')
for h,o,s in es2:
dec=unz(data2[o:o+s])
if dec[:4]!=b'IXUD': continue
# walk null-terminated utf-16le strings in the whole block
toks=re.findall(rb'(?:[\x20-\x7e]\x00){2,}', dec)
toks=[t.decode('utf-16-le') for t in toks]
# pair: text followed by its key
for i in range(len(toks)-1):
if keyrx.match(toks[i+1]) and not keyrx.match(toks[i]):
text[toks[i+1]]=toks[i]
print(f"\nMSG_DEMO text lines loaded: {len(text)}")
# --- render a movie's subtitle transcript ---
def track_for(mv):
k=nh(f"subtitle_{mv}.tbl")
for h,o,s in es:
if h==k: return unz(data[o:o+s])
return None
# index text by integer demo id -> ordered list of lines
bydemo={}
for k,v in text.items():
m=re.match(r'MSG_DEMO_(\d+)_',k)
if m: bydemo.setdefault(int(m.group(1)),[]).append((k,v))
for d in bydemo: bydemo[d].sort()
def transcript(mv):
dec=track_for(mv)
if not dec: return f"[no track for {mv}]"
toks=re.findall(rb'(?:[\x20-\x7e]\x00){2,}', dec)
toks=[t.decode('utf-16-le') for t in toks]
out=[f"=== {mv}.wmv ==="]
# collect all MSG_DEMO ids and timecodes in order; pair positionally
seq=[('id',int(re.match(r'MSG_DEMO_(\d+)$',t).group(1))) if re.match(r'MSG_DEMO_(\d+)$',t)
else ('tc',t) if re.match(r'\d\d:\d\d\.\d\d$',t) else ('x',t) for t in toks]
ids=[v for k,v in seq if k=='id']; tcs=[v for k,v in seq if k=='tc']
for j,demo in enumerate(ids):
tc=tcs[j] if j<len(tcs) else "--:--.--"
lines=[v for _,v in bydemo.get(demo,[])]
out.append(f" {tc} [{demo:3d}] "+(" ".join(lines) if lines else "(no text on disc)"))
return "\n".join(out)
for mv in ["S00A","S16A","hokyu_DS_s02A"]:
print(); print(transcript(mv))