The audio work was three parts and I shipped two. The transcode-fidelity method and the pinned 5.1 downmix landed; the null sink -- the only one that answers "what does the GAME play" -- I deferred to "the next natural rebuild window" and then rebuilt both images four times without doing it. pulseaudio-utils is now in both, with tools/audio-capture wrapping it: a null sink is a real device as far as an application is concerned, so Canary and Godot open it normally and parec records what they emit. This unblocks the decoder's Q8. The cue-to-event bindings are currently a name match against the authors' own identifiers -- a plausible guess, not a measurement -- and capturing what the game plays on a menu move converts them. `audio-capture run` reports the peak level and warns when the capture is silent, because silence is the failure that looks like success: a WAV of exactly the right duration, full of zeroes, because the application opened a different sink. A duration check alone passes it, which is how a confident wrong number gets made.
118 lines
5.3 KiB
Markdown
118 lines
5.3 KiB
Markdown
# Verifying audio without an audio device
|
||
|
||
Neither container has a sound card, so "does it actually play?" cannot be
|
||
answered by listening. It can be answered by measurement, and the two things
|
||
usually meant by that question need different measurements.
|
||
|
||
**Separate them before reaching for a tool:**
|
||
|
||
| question | needs Godot? | needs a device? |
|
||
|---|---|---|
|
||
| Is the transcoded file faithful to the source? | no | no |
|
||
| Does Godot actually route it to an output? | yes | no |
|
||
| What does the *game* play on a menu move? | no (Canary) | a virtual one |
|
||
|
||
## 1. Transcode fidelity — file against file
|
||
|
||
This is the question P4 actually raised, and it needs neither an engine nor a
|
||
device. Decode both, subtract, and measure what is left.
|
||
|
||
```bash
|
||
# Source, for a reference level
|
||
ffmpeg -hide_banner -t 25 -i ADV.wmv \
|
||
-af "aformat=channel_layouts=stereo,astats=measure_perchannel=none" -f null - 2>&1 \
|
||
| grep "RMS level"
|
||
|
||
# The difference signal: source minus transcode
|
||
ffmpeg -hide_banner -t 25 -i ADV.wmv -t 25 -i ADV.ogv -filter_complex \
|
||
"[0:a]aformat=channel_layouts=stereo[a];\
|
||
[1:a]aformat=channel_layouts=stereo,volume=-1[b];\
|
||
[a][b]amix=inputs=2:normalize=0,astats=measure_perchannel=none" -f null - 2>&1 \
|
||
| grep "RMS level"
|
||
```
|
||
|
||
A faithful transcode puts the difference **40 dB or more below** the source.
|
||
|
||
### Three ways this measurement lies
|
||
|
||
Run it wrong and it reports a disaster that is not there. All three of these
|
||
were hit on the first attempt:
|
||
|
||
* **Alignment.** A one-sample offset makes the difference nearly as loud as the
|
||
source. Cross-correlate and compensate *before* subtracting, or the number is
|
||
meaningless. A first run gave source −25.3 dB against difference −34.2 dB —
|
||
only 9 dB down, which looks catastrophic and proves nothing.
|
||
* **Channel layout.** The source and the transcode do not have the same channel
|
||
count. You are not comparing like with like unless both sides are downmixed
|
||
the same way, and `astats` will give you a confident number regardless. See
|
||
[`movie-audio-channels`][mac] for which profile a given movie is in — that is
|
||
a disc fact and lives in the RE corpus, not here.
|
||
* **A file still being written.** `ffprobe` reported the `.ogv` as 33 s against
|
||
the source's 137 s — apparent catastrophic truncation, actually a transcode in
|
||
progress. Check `mtime` and packet count before believing a duration, and
|
||
write to a temp name and rename on completion so a reader cannot see a partial
|
||
file at all.
|
||
|
||
⚠️ **The downmix is an unrecorded decision, and it is not ours to make quietly.**
|
||
Nothing in the manifest says a fold happened or on what weighting; it is whatever
|
||
ffmpeg defaulted to, and that default can change between versions. Centre-channel
|
||
dialogue folds into L/R, so this changes how speech sits against music — an
|
||
aesthetic judgement, not a container detail. Pin it explicitly and record it,
|
||
exactly as MISSION §6 requires of the transcode command itself.
|
||
|
||
[mac]: https://git.mc02.dev/fabi/Syplheed-Reborn/src/branch/main/docs/re/structures/movie-audio-channels.md
|
||
|
||
## 2. Engine routing — Godot writes a WAV instead of a device
|
||
|
||
Godot does not need a sound card to produce audio you can inspect. Put an
|
||
`AudioEffectRecord` on the **Master** bus and it captures the mixed output from
|
||
inside a headless run:
|
||
|
||
```gdscript
|
||
var bus := AudioServer.get_bus_index("Master")
|
||
var rec := AudioEffectRecord.new()
|
||
AudioServer.add_bus_effect(bus, rec)
|
||
rec.set_recording_active(true)
|
||
# ... play the scene ...
|
||
rec.set_recording_active(false)
|
||
rec.get_recording().save_to_wav("user://master.wav")
|
||
```
|
||
|
||
Then feed that WAV through §1 against the source. That closes the loop: it
|
||
proves the asset is right **and** that the engine reached it, which no amount of
|
||
file comparison can show on its own.
|
||
|
||
Confirm the dummy driver is what is actually in use rather than assuming it —
|
||
`AudioServer.get_driver_name()` — and say so in the write-up, because "recorded
|
||
under a dummy driver" is a weaker claim than "heard", and the difference matters.
|
||
|
||
## 3. A virtual device, when something insists on a real one
|
||
|
||
For anything that opens a device rather than a bus — the emulator, most
|
||
obviously — a PulseAudio **null sink** is a real device that records to a file.
|
||
`pulseaudio-utils` is in both images, and `audio-capture` wraps it:
|
||
|
||
```bash
|
||
audio-capture run /tmp/menu.wav -- run-canary # start sink, run, record
|
||
audio-capture start # or drive it by hand
|
||
PULSE_SINK=cap godot --path port
|
||
audio-capture record /tmp/out.wav &
|
||
```
|
||
|
||
This is the route to capturing what the **game** plays — the menu move and
|
||
confirm cues behind HANDOFF Q8 — rather than what we believe it should play.
|
||
Those bindings are currently a name match against the authors' own identifiers;
|
||
a capture turns them into a measurement.
|
||
|
||
⚠️ `audio-capture run` reports the peak level and **warns when the result is
|
||
silent**, because silence is the failure that looks like success: a WAV of
|
||
exactly the right duration, full of zeroes, because the application opened a
|
||
different sink. A duration check alone would pass it.
|
||
|
||
## What none of this establishes
|
||
|
||
That it *sounds right*. Every method here shows correspondence to a source, not
|
||
that the source is the audio the game plays at that moment, and not that levels
|
||
are sane in a mix. A ten-second human listen still answers something no
|
||
measurement above does — so when a result rests on one of these, say which one.
|