check-claims --control plants a revival in docs/port/ and requires exit 1. That the plant lands INSIDE a scanned directory was a property I checked manually, one time, and wrote up -- the exact pattern I had criticised in this same tool one iteration earlier. A fifth case now plants the identical text OUTSIDE the scanned root and requires 0, so the pair asserts the boundary is real: same text, 1 inside and 0 outside. Either half alone is consistent with the tool scanning everything, or nothing. Five cases: clean 0, unmarked 1, marked 0, outside-root 0, empty register 2. audit-kinds has always reported what it found and was never asked whether it can find anything, while its clean runs are cited as evidence that fifteen labels are grounded. --selftest pushes three synthetic rows through the real classifier and reads its verdict: citing nothing must read BARE, a real path ok, a missing path DANGLING. Verified two-directionally -- an extractor stubbed to accept everything returns exit 2. Asserting in check-all. All four submenus are now measured to reset -- LOAD GAME, TUTORIAL and OPTIONS joining EXTRAS -- and the main menu remains the only screen that remembers. Three of the four are not in this export, so no authored value changes. NOT promoted to a rule, deliberately. 'Submenus reset' at 4/4 is better evidence than the 2/2 that made wrap a menu-wide rule, and adopting it would change nothing today because the only submenu this port ships is already measured. What it would do is pre-decide the next screen from a generalisation instead of a measurement -- the trap that nearly let a derived rule overwrite EXTRAS' measured opening item. The guard prints the 4/4 finding beside its per-screen values so the evidence is visible without being load-bearing. MISSION-SELECT-versus-top-item stays open: none of the three separates it, each opens on its own first item, and NEW GAME is untested. Remaining without a harness self-test: verify-transcode-fidelity. Every asserting check passes, 13 of them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF
154 lines
8.3 KiB
Bash
Executable File
154 lines
8.3 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
# Run every check this port has, and say which ones assert.
|
|
#
|
|
# tools/port/check-all
|
|
#
|
|
# There are fourteen tools under `tools/port/` (eleven when this was written --
|
|
# the count is stated because it dates the sentence) and nothing ran them
|
|
# together, so
|
|
# each had to be remembered individually. That is the ninth instance of this
|
|
# port's recurring shape -- something correct, documented and unexercised -- one
|
|
# level up: the checks themselves were the thing nobody was running.
|
|
#
|
|
# ⚠️ It runs the tools that ASSERT. The exploratory ones -- `screen-strip`,
|
|
# `which-focus`, `strip-padding`, `verify-dwell`, `check-capture`,
|
|
# `verify-video-audio` -- produce artifacts for a person to look at and have no
|
|
# verdict to collect. Listing them here as passes would be inventing six.
|
|
set -euo pipefail
|
|
cd "${PROJECT_DIR:-/work}"
|
|
export DISPLAY="${DISPLAY:-:97}"
|
|
OUT="${OUT:-${TMPDIR:-/tmp}/check-all}"; mkdir -p "$OUT"
|
|
BIN="${CARGO_TARGET_DIR:-/sylph-home/port/target-container}/debug/sylpheed-export"
|
|
fail=0
|
|
|
|
step() { # name, expectation, command...
|
|
local name="$1" expect="$2"; shift 2
|
|
local log="$OUT/${name}.log" rc=0
|
|
"$@" >"$log" 2>&1 || rc=$?
|
|
case "$expect" in
|
|
must-pass)
|
|
[ $rc -eq 0 ] && printf ' %-24s ok\n' "$name" \
|
|
|| { printf ' %-24s 🔴 FAILED (rc=%d) -- %s\n' "$name" "$rc" "$log"; fail=1; }
|
|
;;
|
|
report-only)
|
|
printf ' %-24s ran (no verdict -- see below)\n' "$name"
|
|
;;
|
|
esac
|
|
}
|
|
|
|
echo "asserting checks:"
|
|
step format-validator must-pass "$BIN" check
|
|
# The contract lives on a branch this checkout does not merge: HANDOFF on `main`
|
|
# is frozen at 926 lines while the live one is 4 111. Reading 70 unread sections
|
|
# by hand is how two days of deliveries went unread. These are the values that
|
|
# have been reduced to a check; the rest are still read by eye, or not at all.
|
|
step contract-values must-pass tools/port/contract-check
|
|
step contract-control must-pass tools/port/contract-check --control
|
|
# 🔴 The control harness itself is asserted. Every --control run says "each check
|
|
# fails on a perturbed contract"; none of them said "a broken control reports
|
|
# broken". A harness that silently approves a dead check is exactly as useless as
|
|
# a check that silently approves a dead value.
|
|
step control-harness must-pass tools/port/contract-check --selftest
|
|
step modding-rules must-pass tools/port/check-modding
|
|
# Every `kind` in authored/ is a claim about where a value came from, and until
|
|
# 2026-08-30 nothing checked what any of them rested on -- seven were resting on
|
|
# a sibling `why` that argued a different claim.
|
|
step authored-kinds must-pass tools/port/audit-kinds
|
|
# The classifier is asked whether it can tell grounded from ungrounded at all,
|
|
# rather than only what it found. Exit 2 = the harness is broken.
|
|
step kinds-harness must-pass tools/port/audit-kinds --selftest
|
|
# Band levels are alignment-free and carry their own known negative on every run;
|
|
# the difference-signal half of the same tool stays report-only and asserts
|
|
# nothing. See docs/port/DECISIONS.md -- the waveform question is still open.
|
|
step transcode-bands must-pass tools/port/verify-transcode-fidelity
|
|
step capture-controls must-pass tools/port/check-capture-controls
|
|
step menu-audio must-pass env OUT="$OUT/audio" tools/port/verify-menu-audio
|
|
# A stale index is worse than none: it answers "is this already decided?" with a
|
|
# confident no. That is not hypothetical -- see the entry it was built after.
|
|
step decisions-index must-pass tools/port/index-decisions --check
|
|
# A refuted claim asserted outside its correction is a lie the corpus tells a
|
|
# reader who greps for it. Registered claims must carry an explicit `[refuted]`.
|
|
# 🔴 The register check had NO executable control until 2026-08-31 -- every
|
|
# "planted a revival and it failed" in DECISIONS was done by hand, once. Four
|
|
# cases now drive it as a subprocess and read its real exit code, including an
|
|
# EMPTY REGISTER, which used to report clean forever.
|
|
step claims-control must-pass tools/port/check-claims --control
|
|
step refuted-claims must-pass tools/port/check-claims
|
|
echo
|
|
echo "reported, not asserted:"
|
|
step oracle-captures report-only env OUT="$OUT/oracle" tools/port/verify-capture
|
|
sed -n '/^screen /,$p' "$OUT/oracle-captures.log" | sed 's/^/ /'
|
|
# 🔴 `verify-capture` prints and always exits 0. Its own header is right that the
|
|
# numbers are not a target -- the captures carry the game's tone ramp, so RMSE has
|
|
# a floor and driving it lower is fitting the ramp. But "not a target" is not the
|
|
# same as "not a regression detector", and nothing here would notice `title_plate`
|
|
# moving off 0.00 %. Asserting it needs a stored baseline per row, which is a real
|
|
# design decision about what a baseline means when the pose is fitted. NAMED, not
|
|
# quietly skipped.
|
|
|
|
echo
|
|
echo "consistency (expected to differ, for a stated reason):"
|
|
rc=0; env OUT="$OUT/screens" tools/port/verify-screen >"$OUT/verify-screen.log" 2>&1 || rc=$?
|
|
differs=$(grep -c DIFFERS "$OUT/verify-screen.log" || true)
|
|
unexpected=$(grep DIFFERS "$OUT/verify-screen.log" | awk '{print $1}' \
|
|
| grep -vx -e title -e title_jp || true)
|
|
|
|
# 🔴 THE OLD ALLOWANCE WAS FALSE, AND MY FIRST REPLACEMENT REASON WAS ALSO
|
|
# WRONG. Both are recorded because the second error is the more instructive.
|
|
#
|
|
# It said: "the pin is not on main, so this compares two decoder eras". I
|
|
# replaced that with "the eras render identically -- 0 pixels different on three
|
|
# screens". 🔴 **That measurement was void**: the two binaries I compared had the
|
|
# same md5. I built one in a worktree at the pinned tag and one from the
|
|
# workspace, and both commits carry the record-layout fix, so I compared a
|
|
# binary with itself and reported the zero as evidence.
|
|
#
|
|
# Rebuilt properly against `origin/main`, which is the genuinely stale era
|
|
# (`rest t=70 [12 70 80 -]` against the fixed `rest t=12 [0 12 70 80]`):
|
|
#
|
|
# title 0 px main_menu 0 px title_jp 74 507 px
|
|
#
|
|
# ✅ The eras DO change pixels, and `title_jp` is one of the seven bundles where
|
|
# they do -- reproducing the Decoder's figure exactly, under their flags and
|
|
# mine. My "--animated masks it" hypothesis was wrong too.
|
|
#
|
|
# ✅ BUT THE ERA STILL CANNOT EXPLAIN THIS SCRIPT'S ROWS, for a reason I had not
|
|
# established: BOTH SIDES OF THIS COMPARISON ARE THE FIXED ERA. The exporter is
|
|
# pinned to `formats-pin-2026-08-30` and this reference is built from the
|
|
# workspace, and a binary built from each has the SAME md5. There is no era
|
|
# mismatch here to explain anything. Right answer, wrong evidence, and the wrong
|
|
# evidence was a broken experiment.
|
|
#
|
|
# The real reasons are per-screen and already documented:
|
|
# title -- the ptloop SWEEP PHASE residual, max 6 / over3 790, unchanged
|
|
# across every renderer change since P1 (DECISIONS.md).
|
|
# title_jp -- the `--pose=rest` sparkle handling. Adjudicated against the
|
|
# oracle: the port's SHIPPED pose scores r +0.9994 against the
|
|
# game where the reference scores +0.8727, and `--pose=rest` is
|
|
# what this script compares.
|
|
# ⚠️ title_jp is ALSO an era-sensitive bundle, so if this reference is ever
|
|
# built from a different era than the exporter's pin, that row's cause changes
|
|
# and this note stops applying. Check the md5s before trusting it again.
|
|
#
|
|
# So the allowance is now a NAMED SET, not a count with an excuse. A DIFFERS on
|
|
# any other screen fails the run, which a count never could.
|
|
if [ -n "$unexpected" ]; then
|
|
printf ' %-24s 🔴 DIFFERS on %s -- not in the allowed set\n' verify-screen "$(echo $unexpected | tr '\n' ' ')"
|
|
fail=1
|
|
else
|
|
printf ' %-24s %d DIFFERS, both named and explained per screen:\n' verify-screen "$differs"
|
|
printf ' %-24s title = sweep phase; title_jp = rest-pose sparkles (the port is\n' ""
|
|
printf ' %-24s closer to the GAME there than the reference is).\n' ""
|
|
fi
|
|
|
|
# Separately, and unrelated to the rows above: revert to the path dependency when
|
|
# the pin lands. Read from Cargo.toml so it cannot drift out of step again.
|
|
pin=$(sed -n 's/.*tag = "\([^"]*\)".*/\1/p' crates/sylpheed-export/Cargo.toml | head -1)
|
|
if [ -n "$pin" ] && git merge-base --is-ancestor "$pin" origin/main 2>/dev/null; then
|
|
printf ' %-24s ⚠️ %s has landed on main -- revert Cargo.toml to the path dep\n' pin "$pin"
|
|
fi
|
|
|
|
echo
|
|
[ $fail -eq 0 ] && echo "every asserting check passes" || echo "🔴 a check failed"
|
|
exit $fail
|