Files
Sylpheed/docs/port/AUDIO-VERIFICATION.md
Sylpheed port agent f2e08ae31d port: check-capture needed two numbers -- the rate alone passed a 50%-silent file
The Decoder found a blind spot in the bar I shipped last iteration. Raising the
PulseAudio client buffer keeps cutting the gap RATE while total silence bottoms
out and then doubles -- an over-large buffer starves in a few enormous holes
instead of many small ones. Its 500 ms capture scores 1.3 gaps/s, better than a
genuine music bed at 3.3, while being 50% silence. My 20/s bar passed it.

Same shape as the level table that cannot see a duplicated channel: one number,
blind to the failure next door.

I did not set a bar on their numbers, because I do not hold those files and the
last two bars in this tool were wrong precisely from being invented. Instead I
built a control in that regime -- `bigholes`, a real bed with 350 ms holes
punched in -- and set the rule from four controls I can run:

  real music+SFX bed          1.1% silence,  3.3 gaps/s   PASS
  voice track, mono, pauses  53.2% silence,  0.3 gaps/s   PASS
  bed with 350 ms holes      46.3% silence,  3.2 gaps/s   FAIL
  the starved capture        35.6% silence, 30.9 gaps/s   FAIL

Rate alone cannot separate rows 2 and 3; silence alone cannot separate 1 and 3.
The pair does: fail when >=10% is silent on every channel AND there is at least
one gap per second. Real audio is either mostly not silent, or silent in a few
long stretches -- not both at once.

AND THE REGIME IT STILL CANNOT JUDGE IS PRINTED RATHER THAN PASSED. High silence
with very few gaps is what a real voice track looks like and what an
over-buffered capture looks like; nothing here separates them, so the tool says
UNJUDGED and tells the reader to check against a known source. Inventing a bar
for a regime with no control in it is how the previous two bars came to be wrong.

A CONTROL THAT DOES NOT EXECUTE IS NOT A CONTROL: the tool returned immediately
for single-channel input, so the mono voice track -- one of the four controls --
was never run through the check it was meant to control. Mono now skips only the
duplicate test.

Also recorded: the Decoder has withdrawn "the monitor-sink route cannot be fixed
by configuration". A ~200 ms client buffer is worth a retry BEFORE anyone spends
a session on a Canary rebuild.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF
2026-08-29 16:40:24 +00:00

18 KiB
Raw Blame History

Verifying audio without an audio device

Neither container has a sound card, so "does it actually play?" cannot be answered by listening. It can be answered by measurement, and the two things usually meant by that question need different measurements.

Separate them before reaching for a tool:

question needs Godot? needs a device?
Is the transcoded file faithful to the source? no no
Does Godot actually route it to an output? yes no
What does the game play on a menu move? no (Canary) a virtual one

1. Transcode fidelity — file against file

This is the question P4 actually raised, and it needs neither an engine nor a device. Decode both, subtract, and measure what is left.

# Source, for a reference level
ffmpeg -hide_banner -t 25 -i ADV.wmv \
  -af "aformat=channel_layouts=stereo,astats=measure_perchannel=none" -f null - 2>&1 \
  | grep "RMS level"

# The difference signal: source minus transcode
ffmpeg -hide_banner -t 25 -i ADV.wmv -t 25 -i ADV.ogv -filter_complex \
  "[0:a]aformat=channel_layouts=stereo[a];\
   [1:a]aformat=channel_layouts=stereo,volume=-1[b];\
   [a][b]amix=inputs=2:normalize=0,astats=measure_perchannel=none" -f null - 2>&1 \
  | grep "RMS level"

A faithful transcode puts the difference 40 dB or more below the source.

Three ways this measurement lies

Run it wrong and it reports a disaster that is not there. All three of these were hit on the first attempt:

  • Alignment. A one-sample offset makes the difference nearly as loud as the source. Cross-correlate and compensate before subtracting, or the number is meaningless. A first run gave source −25.3 dB against difference −34.2 dB — only 9 dB down, which looks catastrophic and proves nothing.
  • Channel layout. The source and the transcode do not have the same channel count. You are not comparing like with like unless both sides are downmixed the same way, and astats will give you a confident number regardless. See movie-audio-channels for which profile a given movie is in — that is a disc fact and lives in the RE corpus, not here.
  • A file still being written. ffprobe reported the .ogv as 33 s against the source's 137 s — apparent catastrophic truncation, actually a transcode in progress. Check mtime and packet count before believing a duration, and write to a temp name and rename on completion so a reader cannot see a partial file at all.

⚠️ The downmix is an unrecorded decision, and it is not ours to make quietly. Nothing in the manifest says a fold happened or on what weighting; it is whatever ffmpeg defaulted to, and that default can change between versions. Centre-channel dialogue folds into L/R, so this changes how speech sits against music — an aesthetic judgement, not a container detail. Pin it explicitly and record it, exactly as MISSION §6 requires of the transcode command itself.

2. Engine routing — Godot writes a WAV instead of a device

Godot does not need a sound card to produce audio you can inspect. Put an AudioEffectRecord on the Master bus and it captures the mixed output from inside a headless run:

var bus := AudioServer.get_bus_index("Master")
var rec := AudioEffectRecord.new()
AudioServer.add_bus_effect(bus, rec)
rec.set_recording_active(true)
# ... play the scene ...
rec.set_recording_active(false)
rec.get_recording().save_to_wav("user://master.wav")

This is implemented. godot --path port -- --menu … --audio=/tmp/p6.wav installs the effect, records for the whole run, and saves on exit — in _exit_tree rather than beside each quit(), because there are eight of those and the one that would get missed is an error path, i.e. exactly the run whose audio somebody wants to look at. The run prints the driver name beside the file it wrote.

Then feed that WAV through §1 against the source. That closes the loop: it proves the asset is right and that the engine reached it, which no amount of file comparison can show on its own.

Confirm the dummy driver is what is actually in use rather than assuming it — AudioServer.get_driver_name() — and say so in the write-up, because "recorded under a dummy driver" is a weaker claim than "heard", and the difference matters.

3. A virtual device, when something insists on a real one

For anything that opens a device rather than a bus — the emulator, most obviously — a PulseAudio null sink is a real device that records to a file. pulseaudio-utils is in both images, and audio-capture wraps it:

audio-capture run /tmp/menu.wav -- run-canary        # start sink, run, record
audio-capture start                                  # or drive it by hand
PULSE_SINK=cap godot --path port
audio-capture record /tmp/out.wav &

This is the route to capturing what the game plays — the menu move and confirm cues behind HANDOFF Q8 — rather than what we believe it should play. Those bindings are currently a name match against the authors' own identifiers; a capture turns them into a measurement.

⚠️ audio-capture run reports the peak level and warns when the result is silent, because silence is the failure that looks like success: a WAV of exactly the right duration, full of zeroes, because the application opened a different sink. A duration check alone would pass it.

5. A multichannel capture must pass a provenance check BEFORE it is analysed

tools/port/check-capture FILE.wav — run it first, every time.

⚠️ This section exists because a capture of the game's own 6-channel output was analysed at length and the file was corrupt. It got three controls, a drift test and a written-up negative, and every one of those was sound; none of them could see that channels were missing, because the corruption was upstream of everything they tested.

PulseAudio was remapping between two mismatched channel maps, and a 6-channel remap silently drops and duplicates. The Decoder proved it with a control that needs no emulator and no disc — six channels each carrying a different tone, through the same sink and the same parec invocation (docs/re/audio-capture-channel-map-trap.md):

ch played recorded
0 400 400
1 800 3200
2 200 200
3 1600 800
4 3200 800
5 6400 200

Two source channels were gone entirely and two were duplicates. Setting the sink's channel_map to the guest's own (FL,FR,FC,LFE,RL,RR) and passing the same map to parec returns all six.

The signature is an exact duplicate pair, and only a hash finds it

Duration is right. Channel count is right. Corked: no. There is no error anywhere, and the per-channel levels look entirely reasonable — which is the whole difficulty. In the tool's own known-bad control, all six channels report a peak of −18.063656 dB, identical to six decimals, while containing three duplicate pairs. A level check cannot see this. Hashing each channel can.

Two channels of a real surround mix are never byte-identical over tens of seconds. On the corrupt game capture the tool reports:

ch2  peak -4.466272    ba497de78217c438a3e430c5ef6b951b
ch5  peak -4.466272    ba497de78217c438a3e430c5ef6b951b
🔴 ch2 and ch5 are BYTE-IDENTICAL

⚠️ It is a necessary check, not a sufficient one. Passing says the file has no duplicated channels. It says nothing about whether the right thing was recorded — that is what §1's correlation against a known source is for, and a capture should survive both before anything is concluded from it.

Two more conditions, learned the same way

  • Start the recorder before the process you are capturing, so t = 0 precedes it and the window certainly contains the moment of interest.
  • Log what was on screen, with timestamps keyed to the recording's own clock. A capture that matches nothing is then diagnosable rather than ambiguous; the corrupt one could not be told apart from "recorded the wrong phase of the boot" by any amount of analysis at this end.

And the failure this page already warns about, in a second costume: run-canary is silent twice over — SDL_AUDIODRIVER=dummy and --mute=true. Fix only the first and Canary attaches a healthy 6-channel stream at 100 % volume, reports Corked: no, and emits a 19 MB WAV of zeroes.

7. A capture can be starved — right duration, holes punched through it

check-capture tests this too, and it is the second way a recording looks perfect and carries nothing.

A monitor sink advances at wall-clock rate and substitutes silence whenever the producer is late. An emulator running below real time therefore yields a file of exactly the right duration, the right channel count, no duplicated channels — chopped into fragments with holes between them, thousands of times over.

Measured independently on the capture that prompted this (the Decoder's numbers on the untruncated original in brackets):

frames silent on all six channels 35.6 % [39.3 %]
alternating runs 10 482 [10 595]
median burst / gap 13.5 ms / 3.9 ms [13.6 / 3.9]
period 17.4 ms → 57 Hz [≈17.5 ms → 57 Hz]

⚠️ This destroys envelope correlation by construction. What dominates the envelope of such a file is the dropout schedule, not the content — so §6's method was working correctly on a file that could not carry the signal, and the negative it produced said nothing about the game.

Two thresholds I invented were wrong, and the controls caught both

  1. Counting exact-zero frames. Real audio crosses zero constantly, so a clean voice track scored 5 947 "gaps" of median 0.0 ms and was called starved. A gap is a run, not a sample: only runs of ≥ 1 ms count.

  2. Gap count and median length. A genuine music-and-effects bed has 454 gaps at a median of 1.4 ms — quiet 16-bit passages really are zero for milliseconds — so neither statistic separates it from a starved file.

  3. 🔴 The gap RATE alone. This one shipped, and the Decoder found it: raising the client buffer keeps cutting the rate while total silence bottoms out and then doubles, because an over-large buffer starves in a few enormous holes instead of many small ones. Its PULSE_LATENCY_MSEC=500 capture scores 1.3 gaps/s — better than a genuine music bed at 3.3 — while being 50 % silence, and a 20/s bar passed it.

It takes two numbers, because either one alone is blind to the failure next door — the same shape as a level table that cannot see a duplicated channel. Reproduced on a file held here (bigholes: a real bed with 350 ms holes punched into it) so the regime is controlled rather than quoted:

control all-channel silence gaps/s verdict
real music+SFX bed 1.1 % 3.3 PASS
voice track, mono, real pauses 53.2 % 0.3 PASS
bed with 350 ms holes 46.3 % 3.2 FAIL
the starved capture 35.6 % 30.9 FAIL

Rate alone cannot separate rows 2 and 3; silence alone cannot separate rows 1 and 3. The pair does: fail when ≥ 10 % of the file is silent on every channel and there is at least 1 gap per second. Real audio is either mostly not silent, or silent in a few long stretches — not both at once.

⚠️ The regime this tool cannot judge, and says so

High silence with very few gaps is what a real voice track looks like (53.2 % in 0.3 gaps/s) and also what an over-buffered capture looks like. No statistic here separates them. The tool prints UNJUDGED and tells you to check the file against a known source rather than passing it silently — because inventing a bar for a regime with no control in it is how the two bars above came to be wrong.

⚠️ A control that does not execute is not a control. An earlier version returned immediately for a single-channel file, so the mono voice track — one of the four controls — was never actually run through the check it was meant to control. Mono now skips only the duplicate test.

🟡 The monitor-sink route may be fixable after all — retry before rebuilding

An earlier version of this section said the route "cannot be fixed by configuration". Withdrawn. That inferred from the holes that the guest runs below real time, without testing the alternative: the client buffer is simply tiny. Xenia asks SDL for 256 samples — 5.33 ms at 6 ch — against a stock daemon.conf with no fragment tuning.

client buffer silence gaps/s
Xenia default (~5.3 ms) 39.3 % 30.5
PULSE_LATENCY_MSEC=200 15.6 % 3.5
PULSE_LATENCY_MSEC=500 50.1 % 1.3

⚠️ Not clean, and not like-for-like — 88 s against 347 s, and the short run covers the splash logos where silence is real. But the capture route deserves a retry at ~200 ms before anyone spends a session on a Canary rebuild.

The tap, if configuration is not enough

parec reads a monitor that advances at wall-clock rate and substitutes silence, so every moment the emulator runs below real time is a hole, and the timebase is warped non-uniformly — deleting the silences compresses time unevenly rather than repairing it. The route that would work is an internal tap at SDLAudioDriver::SubmitFrame, which sees every frame the guest produces in guest order with no wall clock in the loop.

⚠️ That needs a Canary rebuild, and the Decoder has costed it: build-canary targets a source root that does not exist in that container, the warm build tree is configured against the same missing path, so any change is a full reconfigure plus a full compile on a box with ~700 MB free and a history of parallel builds OOM-killing the host. A whole session for one probe — the human's call, not an agent's.

And a header that never got patched

A streaming writer leaves data declaring 0 bytes. check-capture says so and tells you the duration is unverified — which is not pedantry: the file shared here was copied while it was still being written, and the provenance claim that came with it was wrong about both its length and what it contained.

6. Finding one component inside a mix — and why §1's method cannot

🔴 This section begins with a retraction. Two captures of the game's own output were analysed with sliding envelope cross-correlation and declared not to contain the intro's audio. The instrument was never controlled for the actual task, and when it finally was, it failed:

Can it find the movie's bed inside a synthetic mix of that bed plus the three voice streams? r = 0.415 — below the r > 0.8 bar those negatives were judged against.

The first negative happened to be right (the file was independently proved corrupt by a tone control). It was right by luck, and the reasoning behind it was not supported. A filter that fails its own known-positive is dead, not tuneable.

What was wrong: the threshold, not the idea

r > 0.8 was calibrated on clean-against-clean comparisons, where it is correct — a transcode against its source scores 1.000. A component inside a mix can never score that, because everything else in the mix is uncorrelated noise from the component's point of view. Judging one task by the other's bar guarantees a false negative.

Judge on the LAG and the MARGIN instead. A real match lands at the right lag with a clear gap to the runner-up; a false one is a plateau. And band-limit first, so the component you are hunting dominates what you measure.

The calibration, on a known-present and a known-absent pair

Both bands, both directions, envelope at 0.1 s, minimum 60 s overlap:

hunting band against r lag margin
the movie bed 40–180 Hz mix containing it 0.663 0.0 s ✓ +0.111
the movie bed 40–180 Hz voice-only mix 0.262 −31.9 s ✗ +0.005
voice stream 2 300–3000 Hz mix containing it 0.810 0.0 s ✓ +0.248
voice stream 2 300–3000 Hz the bed alone 0.358 −58.4 s ✗ +0.005

A 20–50× separation in the margin, and the lag is right or absurd. That is a decision rule set by controls rather than by tuning until the data agreed — which is the distinction that matters, and the one the first version of this method skipped.

⚠️ Reach. The known-positive is a synthetic mix at equal gains. A real game mix weights its components differently, so this bounds the method rather than modelling the real case exactly. It is enough to separate present from absent; it is not a level measurement.

What none of this establishes

That it sounds right. Every method here shows correspondence to a source, not that the source is the audio the game plays at that moment, and not that levels are sane in a mix. A ten-second human listen still answers something no measurement above does — so when a result rests on one of these, say which one.

4. What the exporter checks, so nobody has to remember to

sylpheed-export measures peak level and duration of every audio file it writes and records both in manifest.json; sylpheed-export check refuses a tree whose peak is ≤ −90 dBFS (silent) or ≥ 0 dBFS (clipping).

Those are content checks in a format validator on purpose. Silence is the failure this page opens by naming — right duration, right channel count, right size, full of zeroes — and every structural check passes it. Clipping is the other one, and the BGM can produce it, because a music bank is two stems summed at unity gain (HANDOFF Q10).

⚠️ Neither says the audio is the right audio. docs/port/BLOCKED.md says which bindings are measured and which are still authored, and no measurement on this page can move a row there.