The last open question was whether the leading XMA1-mono region duplicates the RIFF sub-wave, which would make the earlier totals double-count. It does not. Decoding both parts of every bank to mono PCM and measuring energy: bank leading secs / RMS riff secs / RMS VOICE_D_450 0.49 / 158 2.82 / 9898 VOICE_D_451 0.01 / 0 1.58 / 9128 VOICE_D_452 0.31 / 301 2.18 / 9061 VOICE_D_453 2.12 / 9770 0.14 / 14462 VOICE_D_454 3.07 / 10428 0.43 / 11639 Two shapes, and no bank holds the same content twice. In 450/451/452 the leading region is silence or near-silence (RMS 0-301 against ~9000 for speech) and the RIFF holds the line. In 453/454 the leading region holds the line and the RIFF is a short loud tail fragment. Sequential segments of one clip, so the totals stand and with them the 48 kHz fit. This also closes the mystery that started the whole thread. The corpus recorded 450 = 2.8 s, 451 = 1.6 s, 452 = 2.2 s as plausible but 453 = 0.14 s and 454 = 0.43 s as "far too short". The decoder skips everything before the first RIFF: for the first three that discards only silence, so they looked fine; for the last two it discards the line itself and leaves the trailing fragment. One rule, two outcomes, depending on which segment holds the speech. The fix is now well-posed in a way the withdrawn attempt was not: emit the leading region only when it carries signal. That also avoids the 1524-bank blast radius that sank the earlier version. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PMRJjbxLqZtsb5Vb7KunPE