The Decoder diagnosed take 2 as a STARVED capture and I verified it here rather
than take it on trust: 35.6% of frames silent on all six channels, 10482
alternating runs, median burst 13.5 ms and gap 3.9 ms, a 17.4 ms period at 57 Hz.
Their untruncated original reads 39.3% and 10595 runs; the difference is exactly
the truncation and every other number agrees.
So my rebuilt correlator was working correctly on a file that could not carry the
signal. The alarming reading it produced -- that the game may not play the .wmv's
WMA track, so ADV.ogv's audio has been wrong since P4 -- is NOT SUPPORTED by this
capture and is not refuted either. Withdrawn as a concern arising from evidence,
with nothing changed in either direction. It was the most expensive-to-act-on
hypothesis in the port and it came from a file that could not speak to it.
THE REAL DEFECT WAS MINE: `check-capture` tested only for duplicated channels, so
it cleared a recording that was 36% holes. A provenance check that passes the
artefact it was built in response to is not a check.
It now measures starvation, and TWO THRESHOLDS I INVENTED WERE BOTH WRONG:
counting exact-zero frames -- real audio crosses zero constantly, so a clean
voice track scored 5947 "gaps" of median 0.0 ms and was called starved. A gap
is a RUN, not a sample; only runs over 1 ms count.
gap count and median length -- a genuine music bed shows 454 gaps at a median
of 1.4 ms, because quiet 16-bit passages really are zero for milliseconds.
What separates them is the RATE: 32.9 gaps/s starved, 3.3 for a real bed, 0.03
for a voice track that is 53% pauses. Bar at 20/s, derived from those controls
rather than chosen and then justified. Controlled both directions: real stereo
bed PASS, six distinct tones PASS, starved capture FAIL. It also reports a `data`
chunk declaring 0 bytes -- what a file copied mid-write looks like, which is what
happened.
VOICE CHANNEL ROLES ARE NOT OBTAINABLE THIS SESSION. Both routes closed: the
monitor sink is starved by construction, and the internal tap at
SDLAudioDriver::SubmitFrame needs a Canary rebuild the Decoder has costed at a
whole session. That is the human's call, not an agent's. The port keeps authoring
with the known recorded.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF
324 lines
15 KiB
Markdown
324 lines
15 KiB
Markdown
# Verifying audio without an audio device
|
||
|
||
Neither container has a sound card, so "does it actually play?" cannot be
|
||
answered by listening. It can be answered by measurement, and the two things
|
||
usually meant by that question need different measurements.
|
||
|
||
**Separate them before reaching for a tool:**
|
||
|
||
| question | needs Godot? | needs a device? |
|
||
|---|---|---|
|
||
| Is the transcoded file faithful to the source? | no | no |
|
||
| Does Godot actually route it to an output? | yes | no |
|
||
| What does the *game* play on a menu move? | no (Canary) | a virtual one |
|
||
|
||
## 1. Transcode fidelity — file against file
|
||
|
||
This is the question P4 actually raised, and it needs neither an engine nor a
|
||
device. Decode both, subtract, and measure what is left.
|
||
|
||
```bash
|
||
# Source, for a reference level
|
||
ffmpeg -hide_banner -t 25 -i ADV.wmv \
|
||
-af "aformat=channel_layouts=stereo,astats=measure_perchannel=none" -f null - 2>&1 \
|
||
| grep "RMS level"
|
||
|
||
# The difference signal: source minus transcode
|
||
ffmpeg -hide_banner -t 25 -i ADV.wmv -t 25 -i ADV.ogv -filter_complex \
|
||
"[0:a]aformat=channel_layouts=stereo[a];\
|
||
[1:a]aformat=channel_layouts=stereo,volume=-1[b];\
|
||
[a][b]amix=inputs=2:normalize=0,astats=measure_perchannel=none" -f null - 2>&1 \
|
||
| grep "RMS level"
|
||
```
|
||
|
||
A faithful transcode puts the difference **40 dB or more below** the source.
|
||
|
||
### Three ways this measurement lies
|
||
|
||
Run it wrong and it reports a disaster that is not there. All three of these
|
||
were hit on the first attempt:
|
||
|
||
* **Alignment.** A one-sample offset makes the difference nearly as loud as the
|
||
source. Cross-correlate and compensate *before* subtracting, or the number is
|
||
meaningless. A first run gave source −25.3 dB against difference −34.2 dB —
|
||
only 9 dB down, which looks catastrophic and proves nothing.
|
||
* **Channel layout.** The source and the transcode do not have the same channel
|
||
count. You are not comparing like with like unless both sides are downmixed
|
||
the same way, and `astats` will give you a confident number regardless. See
|
||
[`movie-audio-channels`][mac] for which profile a given movie is in — that is
|
||
a disc fact and lives in the RE corpus, not here.
|
||
* **A file still being written.** `ffprobe` reported the `.ogv` as 33 s against
|
||
the source's 137 s — apparent catastrophic truncation, actually a transcode in
|
||
progress. Check `mtime` and packet count before believing a duration, and
|
||
write to a temp name and rename on completion so a reader cannot see a partial
|
||
file at all.
|
||
|
||
⚠️ **The downmix is an unrecorded decision, and it is not ours to make quietly.**
|
||
Nothing in the manifest says a fold happened or on what weighting; it is whatever
|
||
ffmpeg defaulted to, and that default can change between versions. Centre-channel
|
||
dialogue folds into L/R, so this changes how speech sits against music — an
|
||
aesthetic judgement, not a container detail. Pin it explicitly and record it,
|
||
exactly as MISSION §6 requires of the transcode command itself.
|
||
|
||
[mac]: https://git.mc02.dev/fabi/Syplheed-Reborn/src/branch/main/docs/re/structures/movie-audio-channels.md
|
||
|
||
## 2. Engine routing — Godot writes a WAV instead of a device
|
||
|
||
Godot does not need a sound card to produce audio you can inspect. Put an
|
||
`AudioEffectRecord` on the **Master** bus and it captures the mixed output from
|
||
inside a headless run:
|
||
|
||
```gdscript
|
||
var bus := AudioServer.get_bus_index("Master")
|
||
var rec := AudioEffectRecord.new()
|
||
AudioServer.add_bus_effect(bus, rec)
|
||
rec.set_recording_active(true)
|
||
# ... play the scene ...
|
||
rec.set_recording_active(false)
|
||
rec.get_recording().save_to_wav("user://master.wav")
|
||
```
|
||
|
||
**This is implemented.** `godot --path port -- --menu … --audio=/tmp/p6.wav`
|
||
installs the effect, records for the whole run, and saves on exit — in
|
||
`_exit_tree` rather than beside each `quit()`, because there are eight of those
|
||
and the one that would get missed is an error path, i.e. exactly the run whose
|
||
audio somebody wants to look at. The run prints the driver name beside the file
|
||
it wrote.
|
||
|
||
Then feed that WAV through §1 against the source. That closes the loop: it
|
||
proves the asset is right **and** that the engine reached it, which no amount of
|
||
file comparison can show on its own.
|
||
|
||
Confirm the dummy driver is what is actually in use rather than assuming it —
|
||
`AudioServer.get_driver_name()` — and say so in the write-up, because "recorded
|
||
under a dummy driver" is a weaker claim than "heard", and the difference matters.
|
||
|
||
## 3. A virtual device, when something insists on a real one
|
||
|
||
For anything that opens a device rather than a bus — the emulator, most
|
||
obviously — a PulseAudio **null sink** is a real device that records to a file.
|
||
`pulseaudio-utils` is in both images, and `audio-capture` wraps it:
|
||
|
||
```bash
|
||
audio-capture run /tmp/menu.wav -- run-canary # start sink, run, record
|
||
audio-capture start # or drive it by hand
|
||
PULSE_SINK=cap godot --path port
|
||
audio-capture record /tmp/out.wav &
|
||
```
|
||
|
||
This is the route to capturing what the **game** plays — the menu move and
|
||
confirm cues behind HANDOFF Q8 — rather than what we believe it should play.
|
||
Those bindings are currently a name match against the authors' own identifiers;
|
||
a capture turns them into a measurement.
|
||
|
||
⚠️ `audio-capture run` reports the peak level and **warns when the result is
|
||
silent**, because silence is the failure that looks like success: a WAV of
|
||
exactly the right duration, full of zeroes, because the application opened a
|
||
different sink. A duration check alone would pass it.
|
||
|
||
## 5. A multichannel capture must pass a provenance check BEFORE it is analysed
|
||
|
||
`tools/port/check-capture FILE.wav` — run it first, every time.
|
||
|
||
⚠️ **This section exists because a capture of the game's own 6-channel output was
|
||
analysed at length and the file was corrupt.** It got three controls, a
|
||
drift test and a written-up negative, and every one of those was sound; none of
|
||
them could see that channels were missing, because the corruption was upstream of
|
||
everything they tested.
|
||
|
||
**PulseAudio was remapping between two mismatched channel maps, and a 6-channel
|
||
remap silently drops and duplicates.** The Decoder proved it with a control that
|
||
needs no emulator and no disc — six channels each carrying a different tone,
|
||
through the same sink and the same `parec` invocation
|
||
(`docs/re/audio-capture-channel-map-trap.md`):
|
||
|
||
| ch | played | recorded |
|
||
|---|---|---|
|
||
| 0 | 400 | 400 |
|
||
| 1 | 800 | **3200** |
|
||
| 2 | 200 | 200 |
|
||
| 3 | 1600 | **800** |
|
||
| 4 | 3200 | **800** |
|
||
| 5 | 6400 | **200** |
|
||
|
||
**Two source channels were gone entirely** and two were duplicates. Setting the
|
||
sink's `channel_map` to the guest's own (`FL,FR,FC,LFE,RL,RR`) and passing the
|
||
same map to `parec` returns all six.
|
||
|
||
### The signature is an exact duplicate pair, and only a hash finds it
|
||
|
||
Duration is right. Channel count is right. `Corked: no`. There is no error
|
||
anywhere, and the **per-channel levels look entirely reasonable** — which is the
|
||
whole difficulty. In the tool's own known-bad control, all six channels report a
|
||
peak of **−18.063656 dB, identical to six decimals, while containing three
|
||
duplicate pairs.** A level check cannot see this. Hashing each channel can.
|
||
|
||
Two channels of a real surround mix are never byte-identical over tens of
|
||
seconds. On the corrupt game capture the tool reports:
|
||
|
||
```
|
||
ch2 peak -4.466272 ba497de78217c438a3e430c5ef6b951b
|
||
ch5 peak -4.466272 ba497de78217c438a3e430c5ef6b951b
|
||
🔴 ch2 and ch5 are BYTE-IDENTICAL
|
||
```
|
||
|
||
⚠️ **It is a necessary check, not a sufficient one.** Passing says the file has no
|
||
duplicated channels. It says nothing about whether the right thing was recorded —
|
||
that is what §1's correlation against a known source is for, and a capture should
|
||
survive **both** before anything is concluded from it.
|
||
|
||
### Two more conditions, learned the same way
|
||
|
||
* **Start the recorder before the process you are capturing**, so `t = 0`
|
||
precedes it and the window certainly contains the moment of interest.
|
||
* **Log what was on screen, with timestamps keyed to the recording's own clock.**
|
||
A capture that matches nothing is then diagnosable rather than ambiguous; the
|
||
corrupt one could not be told apart from "recorded the wrong phase of the boot"
|
||
by any amount of analysis at this end.
|
||
|
||
And the failure this page already warns about, in a second costume:
|
||
`run-canary` is silent **twice over** — `SDL_AUDIODRIVER=dummy` *and*
|
||
`--mute=true`. Fix only the first and Canary attaches a healthy 6-channel stream
|
||
at 100 % volume, reports `Corked: no`, and emits a 19 MB WAV of zeroes.
|
||
|
||
## 7. A capture can be starved — right duration, holes punched through it
|
||
|
||
`check-capture` tests this too, and it is the second way a recording looks
|
||
perfect and carries nothing.
|
||
|
||
**A monitor sink advances at wall-clock rate and substitutes silence whenever the
|
||
producer is late.** An emulator running below real time therefore yields a file
|
||
of exactly the right duration, the right channel count, no duplicated channels —
|
||
chopped into fragments with holes between them, thousands of times over.
|
||
|
||
Measured independently on the capture that prompted this (the Decoder's numbers
|
||
on the untruncated original in brackets):
|
||
|
||
| | |
|
||
|---|---|
|
||
| frames silent on **all six** channels | **35.6 %** [39.3 %] |
|
||
| alternating runs | **10 482** [10 595] |
|
||
| median burst / gap | **13.5 ms / 3.9 ms** [13.6 / 3.9] |
|
||
| period | **17.4 ms → 57 Hz** [≈17.5 ms → 57 Hz] |
|
||
|
||
⚠️ **This destroys envelope correlation by construction.** What dominates the
|
||
envelope of such a file is the dropout schedule, not the content — so §6's method
|
||
was working correctly on a file that could not carry the signal, and the negative
|
||
it produced said nothing about the game.
|
||
|
||
### Two thresholds I invented were wrong, and the controls caught both
|
||
|
||
1. **Counting exact-zero frames.** Real audio crosses zero constantly, so a clean
|
||
voice track scored **5 947 "gaps" of median 0.0 ms** and was called starved. A
|
||
gap is a **run**, not a sample: only runs of ≥ 1 ms count.
|
||
2. **Gap count and median length.** A genuine music-and-effects bed has **454
|
||
gaps at a median of 1.4 ms** — quiet 16-bit passages really are zero for
|
||
milliseconds — so neither statistic separates it from a starved file.
|
||
|
||
**What does separate them is the RATE**, by an order of magnitude in both
|
||
directions:
|
||
|
||
| | gaps/s | median gap | all-channel silence |
|
||
|---|---|---|---|
|
||
| the starved capture | **32.9** | 3.9 ms | 35.6 % |
|
||
| a real music+SFX bed | **3.3** | 1.4 ms | 1.1 % |
|
||
| a voice track, 53 % pauses | **0.03** | — | — |
|
||
|
||
The bar is **20 gaps/s** — 1.6× below the bad case, 6× above the worst good one,
|
||
and set from those controls rather than chosen and then justified.
|
||
|
||
### 🔴 The monitor-sink route cannot be fixed by configuration
|
||
|
||
`parec` reads a monitor that advances at wall-clock rate and substitutes silence,
|
||
so **every moment the emulator runs below real time is a hole**, and the timebase
|
||
is warped non-uniformly — deleting the silences compresses time unevenly rather
|
||
than repairing it. The route that would work is an **internal tap at
|
||
`SDLAudioDriver::SubmitFrame`**, which sees every frame the guest produces in
|
||
guest order with no wall clock in the loop.
|
||
|
||
⚠️ That needs a Canary rebuild, and the Decoder has costed it: `build-canary`
|
||
targets a source root that does not exist in that container, the warm build tree
|
||
is configured against the same missing path, so any change is a full reconfigure
|
||
plus a full compile on a box with ~700 MB free and a history of parallel builds
|
||
OOM-killing the host. **A whole session for one probe** — the human's call, not
|
||
an agent's.
|
||
|
||
### And a header that never got patched
|
||
|
||
A streaming writer leaves `data` declaring **0 bytes**. `check-capture` says so
|
||
and tells you the duration is unverified — which is not pedantry: the file shared
|
||
here was **copied while it was still being written**, and the provenance claim
|
||
that came with it was wrong about both its length and what it contained.
|
||
|
||
## 6. Finding one component inside a mix — and why §1's method cannot
|
||
|
||
🔴 **This section begins with a retraction.** Two captures of the game's own
|
||
output were analysed with sliding envelope cross-correlation and declared not to
|
||
contain the intro's audio. **The instrument was never controlled for the actual
|
||
task**, and when it finally was, it failed:
|
||
|
||
> Can it find the movie's bed inside a synthetic mix of that bed plus the three
|
||
> voice streams? **r = 0.415** — below the `r > 0.8` bar those negatives were
|
||
> judged against.
|
||
|
||
The first negative happened to be right (the file was independently proved
|
||
corrupt by a tone control). **It was right by luck, and the reasoning behind it
|
||
was not supported.** A filter that fails its own known-positive is dead, not
|
||
tuneable.
|
||
|
||
### What was wrong: the threshold, not the idea
|
||
|
||
`r > 0.8` was calibrated on **clean-against-clean** comparisons, where it is
|
||
correct — a transcode against its source scores 1.000. A *component inside a
|
||
mix* can never score that, because everything else in the mix is uncorrelated
|
||
noise from the component's point of view. Judging one task by the other's bar
|
||
guarantees a false negative.
|
||
|
||
**Judge on the LAG and the MARGIN instead.** A real match lands at the *right*
|
||
lag with a clear gap to the runner-up; a false one is a plateau. And **band-limit
|
||
first**, so the component you are hunting dominates what you measure.
|
||
|
||
### The calibration, on a known-present and a known-absent pair
|
||
|
||
Both bands, both directions, envelope at 0.1 s, minimum 60 s overlap:
|
||
|
||
| hunting | band | against | *r* | lag | **margin** |
|
||
|---|---|---|---|---|---|
|
||
| the movie bed | 40–180 Hz | mix containing it | 0.663 | **0.0 s** ✓ | **+0.111** |
|
||
| the movie bed | 40–180 Hz | voice-only mix | 0.262 | −31.9 s ✗ | +0.005 |
|
||
| voice stream 2 | 300–3000 Hz | mix containing it | 0.810 | **0.0 s** ✓ | **+0.248** |
|
||
| voice stream 2 | 300–3000 Hz | the bed alone | 0.358 | −58.4 s ✗ | +0.005 |
|
||
|
||
**A 20–50× separation in the margin, and the lag is right or absurd.** That is a
|
||
decision rule set by controls rather than by tuning until the data agreed —
|
||
which is the distinction that matters, and the one the first version of this
|
||
method skipped.
|
||
|
||
⚠️ **Reach.** The known-positive is a *synthetic* mix at equal gains. A real game
|
||
mix weights its components differently, so this bounds the method rather than
|
||
modelling the real case exactly. It is enough to separate present from absent; it
|
||
is not a level measurement.
|
||
|
||
## What none of this establishes
|
||
|
||
That it *sounds right*. Every method here shows correspondence to a source, not
|
||
that the source is the audio the game plays at that moment, and not that levels
|
||
are sane in a mix. A ten-second human listen still answers something no
|
||
measurement above does — so when a result rests on one of these, say which one.
|
||
|
||
## 4. What the exporter checks, so nobody has to remember to
|
||
|
||
`sylpheed-export` measures **peak level and duration** of every audio file it
|
||
writes and records both in `manifest.json`; `sylpheed-export check` refuses a
|
||
tree whose peak is ≤ −90 dBFS (silent) or ≥ 0 dBFS (clipping).
|
||
|
||
Those are content checks in a format validator on purpose. Silence is the failure
|
||
this page opens by naming — right duration, right channel count, right size, full
|
||
of zeroes — and every structural check passes it. Clipping is the other one, and
|
||
the BGM can produce it, because a music bank is two stems summed at unity gain
|
||
(HANDOFF Q10).
|
||
|
||
⚠️ Neither says the audio is the **right** audio. `docs/port/BLOCKED.md` says
|
||
which bindings are measured and which are still authored, and no measurement on
|
||
this page can move a row there.
|