Take 2 is a good file: it passes check-capture (I re-ran it rather than cite the Decoder's run), carries a screen log, and was recorded with the sink's channel_map set equal to Canary's own. BEFORE REPORTING A SECOND NEGATIVE I ASKED WHETHER MY METHOD COULD DO THE JOB, by building a synthetic mix -- the bed plus the three voice streams -- and hunting the bed inside it. It failed: r=0.415, against the r>0.8 bar my earlier negatives were judged against. So the instrument that produced "the capture contains no ADV audio" could not have found ADV audio in a mix even when it was certainly there. That conclusion was right -- the Decoder's tone control proved take 1 corrupt independently -- but it was right BY LUCK and I reported it as measurement. The three controls I was pleased with tested that the method finds a clean signal in a clean reference, which was never the task. REBUILT AND CALIBRATED IN BOTH DIRECTIONS. Band-limit so the target dominates, then judge on LAG and MARGIN rather than absolute r -- r>0.8 is correct clean-against-clean and meaningless for a component in a mix. bed, 40-180 Hz in a mix containing it r=0.663 lag 0.0 s margin +0.111 bed, 40-180 Hz against a voice-only mix r=0.262 lag wrong margin +0.005 voice, 300-3000 in a mix containing it r=0.810 lag 0.0 s margin +0.248 voice, 300-3000 against the bed alone r=0.358 lag wrong margin +0.005 A 20-50x separation in the discriminating statistic. Written up as AUDIO-VERIFICATION.md section 6, retraction included. THE NEGATIVE NOW STANDS ON SOMETHING. All six of take 2's channels, against both targets, sit in the known-absent regime: margins 0.000-0.017, lags scattered from -72 to +255 s. Take 2 contains neither the movie's WMA bed nor the cutscene voice. Two captures, differently configured, the second provably free of the channel-map fault, with a screen log saying the movie was on screen, and neither carries either source. Handed back: a capture path still losing the mix, or the guest not emitting these sources during the movie, and only one side of the wall can tell those apart. If it is the second it reaches the port directly -- the export's movie audio comes from the .wmv's WMA track. Also noted: the message gives 253.3 s, the file is 318.539 s. The screen log agrees with the file, so it is a mis-stated number, but a length quoted in a provenance claim should match the artefact. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF
255 lines
12 KiB
Markdown
255 lines
12 KiB
Markdown
# Verifying audio without an audio device
|
||
|
||
Neither container has a sound card, so "does it actually play?" cannot be
|
||
answered by listening. It can be answered by measurement, and the two things
|
||
usually meant by that question need different measurements.
|
||
|
||
**Separate them before reaching for a tool:**
|
||
|
||
| question | needs Godot? | needs a device? |
|
||
|---|---|---|
|
||
| Is the transcoded file faithful to the source? | no | no |
|
||
| Does Godot actually route it to an output? | yes | no |
|
||
| What does the *game* play on a menu move? | no (Canary) | a virtual one |
|
||
|
||
## 1. Transcode fidelity — file against file
|
||
|
||
This is the question P4 actually raised, and it needs neither an engine nor a
|
||
device. Decode both, subtract, and measure what is left.
|
||
|
||
```bash
|
||
# Source, for a reference level
|
||
ffmpeg -hide_banner -t 25 -i ADV.wmv \
|
||
-af "aformat=channel_layouts=stereo,astats=measure_perchannel=none" -f null - 2>&1 \
|
||
| grep "RMS level"
|
||
|
||
# The difference signal: source minus transcode
|
||
ffmpeg -hide_banner -t 25 -i ADV.wmv -t 25 -i ADV.ogv -filter_complex \
|
||
"[0:a]aformat=channel_layouts=stereo[a];\
|
||
[1:a]aformat=channel_layouts=stereo,volume=-1[b];\
|
||
[a][b]amix=inputs=2:normalize=0,astats=measure_perchannel=none" -f null - 2>&1 \
|
||
| grep "RMS level"
|
||
```
|
||
|
||
A faithful transcode puts the difference **40 dB or more below** the source.
|
||
|
||
### Three ways this measurement lies
|
||
|
||
Run it wrong and it reports a disaster that is not there. All three of these
|
||
were hit on the first attempt:
|
||
|
||
* **Alignment.** A one-sample offset makes the difference nearly as loud as the
|
||
source. Cross-correlate and compensate *before* subtracting, or the number is
|
||
meaningless. A first run gave source −25.3 dB against difference −34.2 dB —
|
||
only 9 dB down, which looks catastrophic and proves nothing.
|
||
* **Channel layout.** The source and the transcode do not have the same channel
|
||
count. You are not comparing like with like unless both sides are downmixed
|
||
the same way, and `astats` will give you a confident number regardless. See
|
||
[`movie-audio-channels`][mac] for which profile a given movie is in — that is
|
||
a disc fact and lives in the RE corpus, not here.
|
||
* **A file still being written.** `ffprobe` reported the `.ogv` as 33 s against
|
||
the source's 137 s — apparent catastrophic truncation, actually a transcode in
|
||
progress. Check `mtime` and packet count before believing a duration, and
|
||
write to a temp name and rename on completion so a reader cannot see a partial
|
||
file at all.
|
||
|
||
⚠️ **The downmix is an unrecorded decision, and it is not ours to make quietly.**
|
||
Nothing in the manifest says a fold happened or on what weighting; it is whatever
|
||
ffmpeg defaulted to, and that default can change between versions. Centre-channel
|
||
dialogue folds into L/R, so this changes how speech sits against music — an
|
||
aesthetic judgement, not a container detail. Pin it explicitly and record it,
|
||
exactly as MISSION §6 requires of the transcode command itself.
|
||
|
||
[mac]: https://git.mc02.dev/fabi/Syplheed-Reborn/src/branch/main/docs/re/structures/movie-audio-channels.md
|
||
|
||
## 2. Engine routing — Godot writes a WAV instead of a device
|
||
|
||
Godot does not need a sound card to produce audio you can inspect. Put an
|
||
`AudioEffectRecord` on the **Master** bus and it captures the mixed output from
|
||
inside a headless run:
|
||
|
||
```gdscript
|
||
var bus := AudioServer.get_bus_index("Master")
|
||
var rec := AudioEffectRecord.new()
|
||
AudioServer.add_bus_effect(bus, rec)
|
||
rec.set_recording_active(true)
|
||
# ... play the scene ...
|
||
rec.set_recording_active(false)
|
||
rec.get_recording().save_to_wav("user://master.wav")
|
||
```
|
||
|
||
**This is implemented.** `godot --path port -- --menu … --audio=/tmp/p6.wav`
|
||
installs the effect, records for the whole run, and saves on exit — in
|
||
`_exit_tree` rather than beside each `quit()`, because there are eight of those
|
||
and the one that would get missed is an error path, i.e. exactly the run whose
|
||
audio somebody wants to look at. The run prints the driver name beside the file
|
||
it wrote.
|
||
|
||
Then feed that WAV through §1 against the source. That closes the loop: it
|
||
proves the asset is right **and** that the engine reached it, which no amount of
|
||
file comparison can show on its own.
|
||
|
||
Confirm the dummy driver is what is actually in use rather than assuming it —
|
||
`AudioServer.get_driver_name()` — and say so in the write-up, because "recorded
|
||
under a dummy driver" is a weaker claim than "heard", and the difference matters.
|
||
|
||
## 3. A virtual device, when something insists on a real one
|
||
|
||
For anything that opens a device rather than a bus — the emulator, most
|
||
obviously — a PulseAudio **null sink** is a real device that records to a file.
|
||
`pulseaudio-utils` is in both images, and `audio-capture` wraps it:
|
||
|
||
```bash
|
||
audio-capture run /tmp/menu.wav -- run-canary # start sink, run, record
|
||
audio-capture start # or drive it by hand
|
||
PULSE_SINK=cap godot --path port
|
||
audio-capture record /tmp/out.wav &
|
||
```
|
||
|
||
This is the route to capturing what the **game** plays — the menu move and
|
||
confirm cues behind HANDOFF Q8 — rather than what we believe it should play.
|
||
Those bindings are currently a name match against the authors' own identifiers;
|
||
a capture turns them into a measurement.
|
||
|
||
⚠️ `audio-capture run` reports the peak level and **warns when the result is
|
||
silent**, because silence is the failure that looks like success: a WAV of
|
||
exactly the right duration, full of zeroes, because the application opened a
|
||
different sink. A duration check alone would pass it.
|
||
|
||
## 5. A multichannel capture must pass a provenance check BEFORE it is analysed
|
||
|
||
`tools/port/check-capture FILE.wav` — run it first, every time.
|
||
|
||
⚠️ **This section exists because a capture of the game's own 6-channel output was
|
||
analysed at length and the file was corrupt.** It got three controls, a
|
||
drift test and a written-up negative, and every one of those was sound; none of
|
||
them could see that channels were missing, because the corruption was upstream of
|
||
everything they tested.
|
||
|
||
**PulseAudio was remapping between two mismatched channel maps, and a 6-channel
|
||
remap silently drops and duplicates.** The Decoder proved it with a control that
|
||
needs no emulator and no disc — six channels each carrying a different tone,
|
||
through the same sink and the same `parec` invocation
|
||
(`docs/re/audio-capture-channel-map-trap.md`):
|
||
|
||
| ch | played | recorded |
|
||
|---|---|---|
|
||
| 0 | 400 | 400 |
|
||
| 1 | 800 | **3200** |
|
||
| 2 | 200 | 200 |
|
||
| 3 | 1600 | **800** |
|
||
| 4 | 3200 | **800** |
|
||
| 5 | 6400 | **200** |
|
||
|
||
**Two source channels were gone entirely** and two were duplicates. Setting the
|
||
sink's `channel_map` to the guest's own (`FL,FR,FC,LFE,RL,RR`) and passing the
|
||
same map to `parec` returns all six.
|
||
|
||
### The signature is an exact duplicate pair, and only a hash finds it
|
||
|
||
Duration is right. Channel count is right. `Corked: no`. There is no error
|
||
anywhere, and the **per-channel levels look entirely reasonable** — which is the
|
||
whole difficulty. In the tool's own known-bad control, all six channels report a
|
||
peak of **−18.063656 dB, identical to six decimals, while containing three
|
||
duplicate pairs.** A level check cannot see this. Hashing each channel can.
|
||
|
||
Two channels of a real surround mix are never byte-identical over tens of
|
||
seconds. On the corrupt game capture the tool reports:
|
||
|
||
```
|
||
ch2 peak -4.466272 ba497de78217c438a3e430c5ef6b951b
|
||
ch5 peak -4.466272 ba497de78217c438a3e430c5ef6b951b
|
||
🔴 ch2 and ch5 are BYTE-IDENTICAL
|
||
```
|
||
|
||
⚠️ **It is a necessary check, not a sufficient one.** Passing says the file has no
|
||
duplicated channels. It says nothing about whether the right thing was recorded —
|
||
that is what §1's correlation against a known source is for, and a capture should
|
||
survive **both** before anything is concluded from it.
|
||
|
||
### Two more conditions, learned the same way
|
||
|
||
* **Start the recorder before the process you are capturing**, so `t = 0`
|
||
precedes it and the window certainly contains the moment of interest.
|
||
* **Log what was on screen, with timestamps keyed to the recording's own clock.**
|
||
A capture that matches nothing is then diagnosable rather than ambiguous; the
|
||
corrupt one could not be told apart from "recorded the wrong phase of the boot"
|
||
by any amount of analysis at this end.
|
||
|
||
And the failure this page already warns about, in a second costume:
|
||
`run-canary` is silent **twice over** — `SDL_AUDIODRIVER=dummy` *and*
|
||
`--mute=true`. Fix only the first and Canary attaches a healthy 6-channel stream
|
||
at 100 % volume, reports `Corked: no`, and emits a 19 MB WAV of zeroes.
|
||
|
||
## 6. Finding one component inside a mix — and why §1's method cannot
|
||
|
||
🔴 **This section begins with a retraction.** Two captures of the game's own
|
||
output were analysed with sliding envelope cross-correlation and declared not to
|
||
contain the intro's audio. **The instrument was never controlled for the actual
|
||
task**, and when it finally was, it failed:
|
||
|
||
> Can it find the movie's bed inside a synthetic mix of that bed plus the three
|
||
> voice streams? **r = 0.415** — below the `r > 0.8` bar those negatives were
|
||
> judged against.
|
||
|
||
The first negative happened to be right (the file was independently proved
|
||
corrupt by a tone control). **It was right by luck, and the reasoning behind it
|
||
was not supported.** A filter that fails its own known-positive is dead, not
|
||
tuneable.
|
||
|
||
### What was wrong: the threshold, not the idea
|
||
|
||
`r > 0.8` was calibrated on **clean-against-clean** comparisons, where it is
|
||
correct — a transcode against its source scores 1.000. A *component inside a
|
||
mix* can never score that, because everything else in the mix is uncorrelated
|
||
noise from the component's point of view. Judging one task by the other's bar
|
||
guarantees a false negative.
|
||
|
||
**Judge on the LAG and the MARGIN instead.** A real match lands at the *right*
|
||
lag with a clear gap to the runner-up; a false one is a plateau. And **band-limit
|
||
first**, so the component you are hunting dominates what you measure.
|
||
|
||
### The calibration, on a known-present and a known-absent pair
|
||
|
||
Both bands, both directions, envelope at 0.1 s, minimum 60 s overlap:
|
||
|
||
| hunting | band | against | *r* | lag | **margin** |
|
||
|---|---|---|---|---|---|
|
||
| the movie bed | 40–180 Hz | mix containing it | 0.663 | **0.0 s** ✓ | **+0.111** |
|
||
| the movie bed | 40–180 Hz | voice-only mix | 0.262 | −31.9 s ✗ | +0.005 |
|
||
| voice stream 2 | 300–3000 Hz | mix containing it | 0.810 | **0.0 s** ✓ | **+0.248** |
|
||
| voice stream 2 | 300–3000 Hz | the bed alone | 0.358 | −58.4 s ✗ | +0.005 |
|
||
|
||
**A 20–50× separation in the margin, and the lag is right or absurd.** That is a
|
||
decision rule set by controls rather than by tuning until the data agreed —
|
||
which is the distinction that matters, and the one the first version of this
|
||
method skipped.
|
||
|
||
⚠️ **Reach.** The known-positive is a *synthetic* mix at equal gains. A real game
|
||
mix weights its components differently, so this bounds the method rather than
|
||
modelling the real case exactly. It is enough to separate present from absent; it
|
||
is not a level measurement.
|
||
|
||
## What none of this establishes
|
||
|
||
That it *sounds right*. Every method here shows correspondence to a source, not
|
||
that the source is the audio the game plays at that moment, and not that levels
|
||
are sane in a mix. A ten-second human listen still answers something no
|
||
measurement above does — so when a result rests on one of these, say which one.
|
||
|
||
## 4. What the exporter checks, so nobody has to remember to
|
||
|
||
`sylpheed-export` measures **peak level and duration** of every audio file it
|
||
writes and records both in `manifest.json`; `sylpheed-export check` refuses a
|
||
tree whose peak is ≤ −90 dBFS (silent) or ≥ 0 dBFS (clipping).
|
||
|
||
Those are content checks in a format validator on purpose. Silence is the failure
|
||
this page opens by naming — right duration, right channel count, right size, full
|
||
of zeroes — and every structural check passes it. Clipping is the other one, and
|
||
the BGM can produce it, because a music bank is two stems summed at unity gain
|
||
(HANDOFF Q10).
|
||
|
||
⚠️ Neither says the audio is the **right** audio. `docs/port/BLOCKED.md` says
|
||
which bindings are measured and which are still authored, and no measurement on
|
||
this page can move a row there.
|