72f116a3b407872a8e90107b89bdc17c659e37cf
15 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
72f116a3b4 |
port: Ⓐ was never bound to the pad, and the stick is not an edge
Both found by a human playing the port on a real controller. Both were
invisible to every check this port has, for one reason:
`--script` sends InputEventAction, which BYPASSES the input map.
So the harness asserted every line of code AFTER the map and nothing about the
map itself. Measured on this Godot, not remembered -- the remembered answer was
wrong:
ui_accept key:Enter, key:Kp Enter, key:Space <- no joypad at all
ui_cancel key:Escape <- no joypad at all
ui_up key:Up, JOYBTN:11, JOYAXIS:1- <- d-pad AND left stick
ui_down key:Down, JOYBTN:12, JOYAXIS:1+
Four actions worked on the pad and two did not, which presents as a broken
controller: navigation moved, Ⓐ skipped nothing and opened nothing. Godot
4.7.2 binds no joypad button to ui_accept or ui_cancel.
Gamepad.bind_missing() ADDS the two buttons to the built-in actions rather than
redefining them in project.godot, which would replace the built-ins wholesale
and drop the keyboard bindings silently.
Second defect, same blind spot: an InputEventAction is not an analog axis. The
left stick is bound to axis 1, and an axis is not an edge -- held at deflection
it emits an event per jitter, each reporting the action pressed. That was one
cursor step per jitter ("moves the cursor too fast"). The stick is now latched
to one step per deflection, with hysteresis so a stick resting near the
threshold does not chatter.
AUTHORED, and deliberately the conservative half: whether the game REPEATS a
held direction, and how fast, is an oracle question. One deflection one step
cannot run away and invents no rate. Logged as BLOCKED H1.
tools/port/verify-input asserts the map and the latch, with a control that
removes each check's OWN subject -- its first version inverted all nine
assertions when only two depended on the fixup, and reported seven correct
checks as broken. Three rows say plainly they are not controllable (they assert
Godot's own bindings) and one is a negative carrying a positive control (R4),
rather than faking an inversion for either.
Also logged BLOCKED H2, unguessed: the splash blur/fade-in is more pronounced
in the game than in the port. The port applies no blur at all. Noted there that
the two splashes are the only screens reaching the rest() plateau-less
fallback, which the R1 pass just re-opened in both directions.
|
||
|
|
aeb5ef4daf |
port: the shared-state problem is two gaps, and only one needs a human
The Decoder's correction reframes something I had been filing wrongly for a week. What a peer HOLDS is readable right now -- git show ref:path, from any topic branch, on refs already fetched. What a peer must be TOLD still needs a human merge to main. I had been treating both as blocked on the merge; half never was. The symmetry is exact and unflattering to both of us. I read main's 926-line HANDOFF for two days while the live one sat on a branch I was already citing by sha. They read this port's BLOCKED.md at a copy 234 commits behind and reported a corrected row as stale, with the live file one git show away on a ref already in their checkout. Same gap, opposite directions, one command in both. Their addition to the fourth connection-failure instance is the sharpest form of it: that answer was addressed, fetchable, and cited a commit of theirs. Three affordances and neither of us used them. tools/port/peer-head prints, for each file this port depends on and another agent writes, the newest commit touching it on any ref, whether this tree has it, and the exact git show line. Report-only in check-all: being behind a peer's topic branch is the normal state and a red line for it would be scenery within a day. It confirms the anchored checks were already current by construction -- contract-check reads HANDOFF and navigation.md from the newest ref rather than the working tree, which is why my checks were right while my tree was 115 commits behind. It caught a defect in itself on the first run. PROTOCOL.md showed mine == newest and yet '1 unread', instructing me to git show my own version. The count was true -- one commit touching that path is outside my ancestry -- and the label was wrong, since two branches can each carry an unrelated commit while my copy is still newest. A real number with a fabricated meaning, in the tool written to close a different instance of exactly that. Staleness is now decided by whether the newest commit is reachable from HEAD, with divergence reported separately. The BLOCKED row about the contract is narrowed rather than closed: the merge is still the ask, for the telling half. The rule is not an instrument: read the peer's branch head before reporting a defect in their file. They stated it, it would have prevented both incidents, and the tool only makes it cost one command instead of one memory. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
5be071f9cf |
port: close the last control harness, and two authored values checked against bytes
verify-transcode-fidelity --selftest closes my list. It had three controls running every time -- identity, a 4-pole top-end loss, an unrelated movie -- and none asked whether the measurement itself was live. With an empty band list every comparison reads 0.0 dB: identity passes, the real pair passes, and only the unrelated-movie control fails, reporting exit 1 for a broken instrument. Same shape as the empty register in check-claims, same fix: exit 2. The self-test drives the script as a subprocess over a short window -- normal 0, bands emptied 2. All four tools now assert their own harnesses. Top-item sweep from the DIFFICULTY finding: one site, MenuFlow.initial_focus's buttons[0], already documented as a repair. Every other [0] in the tree is unrelated indexing. Nothing to fix, recorded so the sweep is known to have run. The reset question is settled and it went the way that makes the restraint correct: a submenu resets to its OWN OPENING ITEM, a per-screen default that need not be the first. DIFFICULTY opens on NORMAL, second of four, and returns to NORMAL after a confirmed DOWN and a round trip. So ptbtn11 is right for a reason rather than by coincidence, and buttons[0]-is-a-repair is measured rather than principled. contract-check gains check_reset_target, whose teeth the code bounds honestly: on EXTRAS the named item happens to be first, so agreement is not evidence -- what it guards is a future refactor silently substituting an index. Their refutation attempt on extras/initial_focus was made against the disc rather than against their agreement, and it survives: ptbtn11 y282 against 362 and 442. Re-checked from this port's own export, a different reader of the same disc, and the numbers are identical -- extras 282/362/442, main menu 162/242/322/401/482. Which also confirms EXTRAS could never have separated named-item from top-item. Menu focus does not survive a reboot: six fresh boots opened on NEW GAME, three of them following sessions that ended on EXTRAS or OPTIONS. So the authored value is a fresh-start value. The reach is carried verbatim into the why -- every session ended with the emulator KILLED, so this measures 'does not survive a killed session', and a console that remembers across a clean power cycle would not contradict it. Still open and not leaned on: whether the reset target moves once a difficulty has been confirmed; the same SELECT DATA crash prevents testing it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
7ebf5fc8c5 |
port: assert the scan boundary I had hand-verified, and give audit-kinds a self-test
check-claims --control plants a revival in docs/port/ and requires exit 1. That the plant lands INSIDE a scanned directory was a property I checked manually, one time, and wrote up -- the exact pattern I had criticised in this same tool one iteration earlier. A fifth case now plants the identical text OUTSIDE the scanned root and requires 0, so the pair asserts the boundary is real: same text, 1 inside and 0 outside. Either half alone is consistent with the tool scanning everything, or nothing. Five cases: clean 0, unmarked 1, marked 0, outside-root 0, empty register 2. audit-kinds has always reported what it found and was never asked whether it can find anything, while its clean runs are cited as evidence that fifteen labels are grounded. --selftest pushes three synthetic rows through the real classifier and reads its verdict: citing nothing must read BARE, a real path ok, a missing path DANGLING. Verified two-directionally -- an extractor stubbed to accept everything returns exit 2. Asserting in check-all. All four submenus are now measured to reset -- LOAD GAME, TUTORIAL and OPTIONS joining EXTRAS -- and the main menu remains the only screen that remembers. Three of the four are not in this export, so no authored value changes. NOT promoted to a rule, deliberately. 'Submenus reset' at 4/4 is better evidence than the 2/2 that made wrap a menu-wide rule, and adopting it would change nothing today because the only submenu this port ships is already measured. What it would do is pre-decide the next screen from a generalisation instead of a measurement -- the trap that nearly let a derived rule overwrite EXTRAS' measured opening item. The guard prints the 4/4 finding beside its per-screen values so the evidence is visible without being load-bearing. MISSION-SELECT-versus-top-item stays open: none of the three separates it, each opens on its own first item, and NEW GAME is untested. Remaining without a harness self-test: verify-transcode-fidelity. Every asserting check passes, 13 of them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
575e287526 |
port: the register check had no executable control, and an empty register passed forever
check-claims guards the refuted register, the thing both agents lean on when they say a dead claim is not being re-asserted, and it had no control machinery at all. Every 'planted a revival, it failed, removed it, it passed' in DECISIONS was done by hand, once, and never again -- in a repository where two of my own tools carry the line 'a control that does not execute is not a control'. I wrote that about somebody else's tool. The hole the Decoder found in their equivalent was here too. The scan loop runs once per register row; with no rows it runs zero times, fail stays 0, and the script printed 'every refuted claim appears only inside its correction' and exited 0. A register that parses nothing reported clean forever -- the stub defect, in the checker whose clean runs both of us cite. It now exits 2 with 'the harness is broken, not the corpus'. --control executes four cases, each driving this script as a subprocess and reading its real exit code: clean 0, unmarked revival 1, marked revival 0 with no false positive, empty register 2. Asserting in check-all. Two things taken from their build of the same thing rather than invented: the self-test drives the real machinery and reads its actual exit code -- my first --selftest reasoned about what the harness would do, which is the cheaper mistake and the one I made -- and the three-way exit convention, which is what lets 'the corpus is dirty' and 'the checker is broken' be different answers instead of both being nonzero. The plant lands in a real scanned directory, because a control that runs somewhere the tool does not look proves nothing about the tool. Verified two-directionally: pointing the plant at an unscanned path makes the control report itself broken. Still without harness self-tests and filed rather than left looking finished: audit-kinds and verify-transcode-fidelity. Every asserting check passes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
14c3bad7ad |
port: the control harness now asserts itself, and it caught me twice doing it
The gap I named and the Decoder prioritised: every --control run asserts that each check fails on a perturbed contract, and none asserted that a broken control reports broken. That is printing a verdict without asserting it, one level up. A harness that silently approves a dead check is exactly as useless as a check that silently approves a dead value. contract-check --selftest feeds the machinery a stub that cannot fail -- a function that prints 'everything is fine' and asserts nothing, which is precisely the defect I shipped in verify-transcode-fidelity's unconditional return 0 -- and requires the machinery to flag it. Exit codes follow the Decoder's convention: 0 all good, 1 a real check failed, 2 the HARNESS is broken and nothing it reported can be trusted. Asserting in check-all. It caught two defects while being written. The first version checked that the stub left the failure counter at zero and then REASONED that control() would therefore flag it -- arguing where a measurement was available, the error this whole thread has been about, committed inside the tool built to prevent it. Rewritten to push the stub through the real control() loop and read its verdict. It then returned 2 immediately: the stub was flagged, but as 'the control's own anchor is gone' rather than as a dead check, because the src selection anchored anything not in one specific list at the walk document instead of HANDOFF. A real failure for a fabricated reason, which is the confusion ANCHOR SPLIT exists to separate. Not covered and filed rather than left looking finished: check-claims, audit-kinds and verify-transcode-fidelity have controls and no harness self-test. The shape is known and the fix is cheap. Every asserting check passes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
de25787d84 |
port: a capital letter hid a refuted claim; and band levels answer what alignment could not
Three findings, two of them defects in my own checkers. Changing the KIND of quantity answered the P4 fidelity question on the first attempt. Four attempts at sample-exact difference-signal alignment produced four failures and no verdict -- well past the Decoder's rule that two failed attempts at the same measurement are evidence the quantity is wrong, not the parsing. Band energies need no alignment at all: both transcodes match their sources to 0.66 dB worst-case across four bands, while an unrelated movie lands at 19-20 dB. Two populations an order of magnitude apart, so the 1.5 dB tolerance sits between measured values rather than being picked. Asserting in check-all with the known negative on every run, not behind a flag. It also diagnoses the failure it replaced: matching spectra mean same content at same level, so the difference signal's failure is my alignment, now by evidence rather than assumption. The difference path stays report-only. Band agreement cannot tell a faithful transcode from one that kept the spectrum and mangled the waveform -- weaker than P4 wanted, and what I can support. check-claims held 'no loop-point field has been identified' in its register the whole time and matched case-sensitively, so a capital N at the start of a sentence hid a registered dead claim in BLOCKED.md -- the one document whose job is to say what is still open. The correction had reached authored/audio.json and not the blocked list, which is exactly the failure that file's own why warns about. Matching is case-insensitive now and immediately surfaced five more unmarked sites, including a whole DECISIONS section still describing the refuted state. All six fixed: four tokened, two rewritten with the shipped values. Controlled with a planted capitalised revival. And --control caught its own harness: it perturbed only the first occurrence of an anchor, and the Decoder's delivery heading now appears twice, so the check read the untouched duplicate and passed a wrong contract. A perturbation that does not reach every copy makes a check untestable silently. First time a control has failed because of a change in someone else's document rather than my code. Not accepted from the same message: the (A)-skips-a-movie row is NOT stale. It reads (a) ANSWERED, cites Q9, and points at flow.json's skippable: true. Reported back rather than quietly 'fixed' -- marking a live row stale is the error their own message is about. Every asserting check passes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
7c2a47f8be |
port: audit every kind label, and seven rested on a neighbour's argument
tools/port/audit-kinds reports what each in authored/ rests on. Nothing had ever checked them, which is the point -- the disciplines that fail this way are the ones that never visibly failed. Seven of fifteen labels, every goto_name_kind, had no of their own. Four scored ok on the first run because the audit fell back to the parent's , which argues the DESTINATION while the label is about where the NAME came from. That is the same error I was corrected for the previous iteration, one level down: crediting a claim with evidence that does not bear on it. Borrowed evidence is now its own outcome, and all seven carry a why citing HANDOFF Q4's own words and stating that the port never branches on the field. The audit refuted itself twice first. It counted only paths, shas and filenames as citations, so HANDOFF Q1 and PORT-MISSION section 7 read as citing nothing -- four false positives, and an audit that invents defects is worse than none because its false positives are indistinguishable from its true ones until each is opened. It also resolved paths against committed refs only, failing on a citation to the tool being written. Both fixed. It still cannot read a cited page to confirm it says what the why claims, and prints that every run. MEASURED and measured both existed; a consumer comparing == measured misses the other, and a label that fails to match reads as ABSENT rather than wrong. Normalised. Refutation attempt on HANDOFF Q2's map of GP_TITLE. The headline survives and is exactly right: 4 UI states + 2 loading variants + 2 boot splashes = 8 states shipped twice = the 16 entries the archive holds, confirmed against my export's entry map. But the row enumerates six of those eight -- entries 10, 11, 13 and 14, publisher_logo and developer_logos, appear nowhere in it. A reader counting Q2 gets twelve, and this is the row already corrected once for an ordinal-versus- entry error, which is the mistake four unlisted entries feed. The port is unaffected; both splashes are exported, named and verified at RMSE 2.17 and 3.05. Every asserting check passes, audit-kinds included. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
ad80938efc |
port: check the contract's numbers instead of reading 4111 lines of it
HANDOFF on main is 926 lines frozen at 0fd8e69; the live one is 4111 at
|
||
|
|
ad8e18de7c |
port: sweep instructions above descriptions -- the silent class is clean, two loud hits
Their sharpening: a stale instruction manufactures a false confirmation, strictly worse than a stale description that merely misleads. Applied to my instruction surface, the documented invocations in tool and script headers. All fifteen distinct flags across those examples are parsed, so nothing in my headers can produce their failure mode by being inert. But 'parsed' is a proxy and its gap is known -- --shots parses and does nothing on the --boot path -- so I ran two documented examples end to end rather than trusting the grep, and both produce a 1280x720 frame. Two hits, both loud rather than silent: 11 references to tools/verify-capture and tools/verify-screen, paths that do not exist since the tools are under tools/port/ (fixed in 4 files); and check-all claiming eleven tools where there are fourteen (now states both so the sentence dates itself). The distinction worth recording: mine fail loudly, theirs failed silently. A wrong path announces itself; an inert environment variable returns a clean wrong result. Both are stale instructions and only one manufactures evidence. Honest limit: I tested the flag surface plus two examples end to end, not all thirteen documented invocations -- the --boot ones take 156 s each. That is a judgement about cost, not a claim of coverage. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
e11d843e72 |
port: WITHDRAW the 'eras render identically' measurement -- I compared a binary with itself
Last iteration I overturned check-all's allowance on a measurement of 0 pixels between the two decoder eras, and rewrote the tool's reason around it. The two binaries had the same md5: one built in a worktree at formats-pin-2026-08-30 and one from the workspace, and both commits carry the record-layout fix. I compared a binary with itself and reported the zero as evidence. The 508-line diff I cited was real and irrelevant -- it does not straddle the fix. Done properly against origin/main, verified stale by the Decoder's own control (rest t=70 vs rest t=12) and by differing md5s: title 0 px, main_menu 0 px, title_jp 74507 px -- reproducing their figure exactly, under their flags and mine. My second hypothesis, that --animated masked it, was also wrong. What survives: the era still cannot explain this script's rows, for a fact I had not established -- both sides of the comparison are the FIXED era, since a binary built from the pin and one from the workspace have the same md5. Right answer, wrong evidence. The note now carries its condition: title_jp is era-sensitive, so if the reference is ever built from a different era than the pin, that row's cause changes. Twice now a correct conclusion has come through a broken experiment, and both times the tell was two things that should differ producing identical output. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
ed14722996 |
port: check-all excused two failing rows with a measurably false reason
The suite reported '2 DIFFERS, allowed: the pin is not on main, so this compares two decoder eras', and I had quoted that for several iterations without testing it. Built sylpheed-cli at formats-pin-2026-08-30 and at workspace HEAD and rendered through both: title, title_jp and main_menu come out 0 pixels different, despite 508 lines of difference in ui_layout.rs. The eras are not the cause, and the allowance was excusing a real signal with a wrong explanation. A second defect in the same eight lines: the expiry tested formats-pin-2026-08-29d while Cargo.toml pins formats-pin-2026-08-30, so it would have expired on a tag this tree does not use. The real reasons are per-screen and already documented: title is the ptloop sweep phase residual, title_jp is the --pose=rest sparkle handling -- where the port's shipped pose scores +0.9994 against the game to the reference's +0.8727, so the port is closer to the game on the row the script calls a disagreement. Replaced with a named set: title and title_jp by name, any other DIFFERS fails. A count cannot notice a different screen drifting while the total stays at two. Controlled both directions -- passes on the known pair, fails on main_menu or extras. The pin reminder now reads the tag out of Cargo.toml so it cannot drift. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
769ac9a2c7 |
port: a refuted-claim register, enforced by check-all
The Decoder's audit of their own corpus found four refuted claims standing -- including one they had corrected to me, agreed with, and written a METHOD entry about, without landing it for a full iteration. A hand audit finds what is there on the day it runs; it does not stop the next one. check-claims is a register: every occurrence of a refuted claim must carry an explicit [refuted] sentinel within 400 characters. It found four more unmarked occurrences than my manual pass had, including one in authored/audio.json. The marker is a sentinel rather than a keyword because the first version's every failure was a quotation inside a correction whose wording lacked the keyword. The temptation was to widen the window until they passed -- tuning a threshold until the answer comes out right, in the tool built to catch that. 21 quotations marked by hand; proved it fails by removing one. Also fixes the Decoder's other finding in my corpus: BLOCKED's voice row had a struck heading with three sentences below still asserting in the present tense. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
19c0aa89f1 |
port: index DECISIONS.md -- it already answered last iteration's question
Last iteration I filed title and title_jp's disagreement with sylpheed-cli as mechanism-unknown, to the Decoder as well as here. Both were already explained in this file, under headings that name the two screens. Checked rather than assumed. title: still ties on 0x8083, 0x80a0 and 0x8010, and the export declares paint_order_ties unresolved; the old entry's 904 px in the glow band matches my 790 px at the same place, same 4-6/255 magnitude. title_jp: the 'only non-integer scale' claim finds 26 keyframes export-wide, but exactly ONE element visible at rest -- ptlogo_eff2 at 125% -- which is the pose verify-screen uses. It survives narrowly. The failure is navigability: 6502 lines, 111 sections, no index, so 'has this been decided?' had no cheap answer and re-deriving it looked like diligence. index-decisions generates the contents; check-all runs --check. No line numbers (the first version was a fixpoint that failed its own check, and appends would invalidate them all), and checked, because a stale index answers 'already decided?' with a confident no. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
49b49668c3 |
port: add check-all; verify-screen ignored its own statistic; 'six expected DIFFERS' was wrong
Eleven tools and nothing ran them together -- the ninth instance of correct, documented and unexercised, one level up. check-all runs the four that assert, reports the oracle table, and gives verify-screen an allowance that EXPIRES when the pin lands rather than standing forever. All eleven exercised first; none had rotted. verify-screen computed over3 because 'a single max cannot tell 2 pixels from 25 444' and then decided the verdict on max alone: main_menu (max 4, over3 0) read DIFFERS while extras (max 3, over3 0) read OK. The bar is unchanged; a frame with no pixel over it now gets its own ROUNDING verdict. And corrects a claim I have given the Decoder more than once. The real count was ten, now eight: six forced-backdrop, two rounding, and TWO UNEXPLAINED -- title at 790 px and title_jp at 20498, neither carrying a forced element. My leaf hypothesis is refuted: emptying draw_leaf_for changes the numbers not at all. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |