The 60 tokens absent from BOTH voice languages, itemised:
45 VOICE_E_ family -- 44 numeric [0..43] plus the lettered VOICE_E_012B
13 VOICE_C_ at 421, 423-426, 430, 432, 447-450, 470, 471 -- INSIDE the
listed range [0..489], so interior gaps rather than a truncated tail
2 VOICE_D_182 and _183, adjacent
These lines were written and captioned; only the audio is missing. Their
caption keys resolve to real text in the IXUD blocks:
MSG_VOICE_C_355_000_00 "What are you doing? Quit wasting..."
MSG_VOICE_C_367_000_00 "The final defense weapon is..."
MSG_VOICE_C_347_000_00 (Japanese)
MSG_VOICE_D_152_000_00 (Japanese)
MSG_VOICE_E_044_000_00 (Japanese)
Three of the five sampled are still Japanese INSIDE the English pak --
captioned but never translated, matching the untranslated entries already
noted for the localised-text container.
Separately, a trap worth its own heading: within one message page the voice
bank token and the caption keys use DIFFERENT numbering.
Message_106 voice VOICE_C_468 lines MSG_VOICE_C_385_000_00..02
Message_129 voice VOICE_D_182 lines MSG_VOICE_D_152_000_00..02
Message_044 voice VOICE_E_012B ID MSG_VOICE_E_044
Same family letter, different index space. Deriving one id from the other
will silently mis-pair audio with text -- which is the same class of mistake
as the demo-id voice binding this corpus already had to reject in-game.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PMRJjbxLqZtsb5Vb7KunPE