This corrects the previous commit. 文字列 is not a developer's leftover: it is one member of a six-word Shift-JIS TYPE vocabulary, and the records carrying it are a machine-readable schema for the unit datasheet. The whole non-ASCII population on the disc is 6 distinct values out of 99328 - 0 of 3496 record names and 0 of 12173 field names - and all six are type words: 文字列 string 366 uses NS_"文字列" NS_ string 12 整数 / 整数値 integer 66 / 6 浮動小数値 floating-point value 504 浮動小数値[0〜1] float in [0,1] 36 990 type-valued fields. So the reader defect noted last time is real but bounded to these six strings, and name_hash re-encodes Latin-1 byte-for-byte, so hashing was never affected. They sit in 15 records x 6 GP_MAIN_GAME_* paks = 90 instances, i.e. 15 records with ONE user. The names are exactly the unit substructure family, and six carry a literal wildcard: Turret_???, Hatch_???, Bridge_???, Thruster_???, ShieldGenerator_???, Versatile_???, and NS_*. ??? is the numeric-suffix wildcard at record AND field level - Turret_??? is the schema for Turret_000..00N, and inside it CannonFrame_??? / MuzzleFrame_??? stand for the numbered slots. Where a field's type is an enumeration the schema holds an EXAMPLE value instead of a type name: Yes for the five booleans, Vessel for Generic.Type (the 43 Craft + 71 Vessel split), Ship_ for the ID prefix convention. Maneuver is the one fully-typed record, 34 of 34. Every ResistanceTo* and every Color_* channel is declared FLOAT[0..1] - normalised by declaration, matching the sampled values in unit-datasheet-static. Generic.NozzleSpec_??? has its own type NS_"文字列" and NS_* is a record, so the nozzle spec is a nested sub-schema. Control separates schema from data cleanly: the _??? records and NS_* exist ONLY as schema, 6 of 6 instances typed, while the eight real substructure names are typed in 6 instances and untyped in the rest - Generic 6 of 3651, the others 6 of 684 each. Turret_??? carries the game's own typo NomalModel beside DamagedModel. This gives the port an authoritative field-type table: types the disc declares, rather than types inferred from sampled values. New artefact with its regenerator: tools/re-capture/datasheet_schema.py -> docs/re/data/datasheet-schema.txt. All fifteen existing artefacts byte-identical.
57 lines
2.2 KiB
Python
57 lines
2.2 KiB
Python
#!/usr/bin/env python3
|
|
"""The unit datasheet's own schema, as shipped on the disc.
|
|
|
|
Each GP_MAIN_GAME_*.pak carries 15 IDXD records whose field VALUES are Japanese
|
|
type names rather than data -- Shift-JIS, and the only non-ASCII strings on the
|
|
whole disc. Records whose name ends '_???' (and 'NS_*') are wildcards standing
|
|
for the numbered family, e.g. Turret_??? for Turret_000..Turret_00N.
|
|
|
|
Where a field's type is an enumeration the schema holds an EXAMPLE value
|
|
instead of a type name ('Yes' for a boolean, 'Vessel' for Generic.Type).
|
|
|
|
Regenerates docs/re/data/datasheet-schema.txt.
|
|
"""
|
|
import glob, os, sys
|
|
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
|
|
from unit_substructures import pak_entries
|
|
import unitgroup as U
|
|
|
|
# the six type names, held as latin-1 (how unitgroup.py hands strings back)
|
|
TYPES = {
|
|
'\x95\xb6\x8e\x9a\x97\xf1': 'STRING',
|
|
'NS_"\x95\xb6\x8e\x9a\x97\xf1"': 'NS_STRING',
|
|
'\x90\xae\x90\x94': 'INT',
|
|
'\x90\xae\x90\x94\x92l': 'INT_VALUE',
|
|
'\x95\x82\x93\xae\x8f\xac\x90\x94\x92l': 'FLOAT',
|
|
'\x95\x82\x93\xae\x8f\xac\x90\x94\x92l[0\x81`1]': 'FLOAT[0..1]',
|
|
}
|
|
|
|
def main():
|
|
pak = sorted(glob.glob('/work/sylph_extract/**/GP_MAIN_GAME_D.pak', recursive=True))
|
|
if not pak:
|
|
print('disc not mounted; nothing to do'); return
|
|
out = {}
|
|
for h, b in pak_entries(pak[0]):
|
|
if b[:4] != b'IDXD':
|
|
continue
|
|
try:
|
|
recs = U.parse(b)
|
|
except Exception:
|
|
continue
|
|
for r in recs:
|
|
if any(v in TYPES for t, f, v in r['fields']):
|
|
out[r['squadron']] = r['fields']
|
|
print('the unit datasheet schema shipped in GP_MAIN_GAME_*.pak')
|
|
print('%d schema records; identical in all six paks (x6 = one user)' % len(out))
|
|
print()
|
|
for name in sorted(out):
|
|
fields = out[name]
|
|
typed = sum(1 for t, f, v in fields if v in TYPES)
|
|
print('=== %s (%d fields, %d typed) ===' % (name, len(fields), typed))
|
|
for t, f, v in sorted(fields, key=lambda x: (x[1] or '')):
|
|
print(' %-28s %s' % (f, TYPES.get(v, 'example: ' + repr(v))))
|
|
print()
|
|
|
|
if __name__ == '__main__':
|
|
main()
|