movie_subtitle handles MSG_DEMO_*, the cutscene captions. Counting every IXUD
block in GP_MAIN_GAME_E.pak, that is the SMALLEST of eight families:
MSG_ADAN 23236 keys 9801 with text ADAN combat chatter
MSG_RHIN 21196 8509 Rhino squadron
MSG_TCAF 17148 6728 TCAF
MSG_VOICE 13060 6776 in-mission scripted dialogue
MSG_BIRD 14036 5834 Bird squadron
MSG_ADPL 12640 4127 ADAN pilots
MSG_ACRO 4804 2244 Acropolis
MSG_DEMO 1252 560 cutscene captions <- the only one read
total 107372 44579
560 of 44579 text-bearing keys = 1.3%. I report the text-bearing column rather
than raw keys because only 41.5% of keys carry text -- the rest are the empty
line slots this container pads with, and counting those would flatter the
denominator.
MSG_VOICE_* is the family the message tables reference -- the dialogue whose
voice bindings this file now analyses in detail -- and nothing in crates/
parses it. So the corpus knows which bank plays for a line it cannot read.
First step recorded: build_demo_text already pairs a text value with the
MSG_DEMO_<demo>_<page>_<line> key that follows it, and the other seven
families use the same <id>_<page>_<line> shape, so generalising the key parser
is most of the work. With a warning attached: do NOT assume the id spaces
relate, since the voice-bank id and the caption id within one message page are
different numbers.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PMRJjbxLqZtsb5Vb7KunPE