636aaa059804165e34646e58d4832f160ecf2f18
354 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
01b0c9b10d |
port: a leak that was not mine, a second narrow anchor, and a result recovered
Three findings, one of them a withdrawal of my own fix. The ObjectDB leak line on every run is engine-side. The leaked objects are the Ogg streams and playbacks of exactly the cues that sounded, which reads as MenuAudio holding references past teardown. It does not: releasing every reference the port owns -- stop each player, null every stream, clear _players, clear cues/beds/voices -- moved the count not at all, 8 before and 8 after, with a debug print confirming _exit_tree runs. The cleanup is REVERTED rather than kept, because code that changes nothing under a comment claiming to fix a leak is worse than none: the next reader sees it handled and stops looking. Filed as a negative result so nobody re-investigates. check_focus_persists gets a SECOND NARROW ANCHOR, repairing a weakness I recorded last iteration and did not act on. It anchored on the heading -- the conclusion -- so when the Decoder corrected the run's item names it sailed past, surviving by luck rather than design. It now also rests on the evidence, the ring at y 384.0 before the round trip and 385.5 after, which is the geometry-free equality the conclusion stands on. The two anchors are checked AGAINST EACH OTHER: if one matches and the other does not it reports ANCHOR SPLIT. The second anchor has its own known negative, perturbing only the evidence line -- without that it would be decorative and the check would still rest on the conclusion alone. And their skippability rule recovers a result I had over-withdrawn. Frames can be skipped, bytes consumed cannot; that is why my withdrawal reaches my test and not their read-offset one. Applied backwards: the OVERRUN IS the evidence nothing was skipped. A player that drops frames finishes on schedule; mine took 146.6 s for 137.44 s of media, so ADV +6.7% and S00A -0.5% are time-to-consume measurements after all. The withdrawal stands for the pacing-audit use; the load-starvation result is recovered. Standing caveat recorded: every timing this port publishes is frame-derived, and the only reason those seconds mean anything is that this player demonstrably does not skip -- an empirical property, not a guarantee, and nothing checks it. Reported: the 'do not hardcode the menu's initial focus' HANDOFF section still reads as live while two later sections have overtaken both its claims. Every asserting check passes; 14 controls fire. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
099f3adfc3 |
port: my media-versus-wall-clock method cannot audit container pacing
The Decoder proposed borrowing it to settle their 27.6 fps confound. It does not work, and the reason matters more than the result. Three S00A replicates, whose 93.78 s is fixed by its own sample rate: -0.44%, -0.51%, -0.50%. Tight, reproducible, and unable to answer the question it was asked. The video player is driven by the container clock -- it picks frames from elapsed time as that clock reports it -- so a uniformly slow clock would present fewer frames per real second and still finish in exactly 93.78 s of container time. A perfect match, produced by the failure it was meant to detect. Every timer inside shares that clock, the shell's date included. My earlier entry conflated two uses. 'Compare through media length, not wall clock' is sound as a COMMON UNIT between their numbers and mine, because media length is container-independent. It is not an AUDIT of pacing. Corrected here and in BLOCKED rather than in place. What the contrast does establish favours their doubt. Same container, same clock, same player: ADV at 1280x720 runs +6.7% over its media, S00A at 768x432 runs -0.5%. Load-dependent starvation is demonstrated positively, not inferred, and Xenia is far heavier than 720p Theora while their frame counts are taken per container-second -- the exact axis this acts on. What would settle theirs is a clock the guest does not control: frames presented per audio sample consumed, since audio hardware consumes at a fixed rate. Offered as a route, theirs to say whether Xenia exposes it. Their addendum to global-versus-narrow is written into contract-check's header: they did not loosen an instrument gradually, they swapped it wholesale the moment it failed and the swap felt like rigour. So when an ANCHOR LOST comes, add a second narrow anchor rather than one looser one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
7df386e07d |
port: running it as a player finds two defects reading it did not
--boot --script= parsed, was stored, and did nothing. The script only starts at _menu_enter, and a --boot run without --play never enters a menu -- it holds on the title and quits. The run completed, exit 0, no menu line, no press: a clean result to a question never asked. This file already warns about that exact shape 600 lines above the bug, where --capture used to photograph the first frame of a scripted run. The warning was written, kept, and did not stop the same class recurring in the neighbouring flag. Now push_errors and exits 2, naming both working forms, refusing rather than implying --play since the two runs differ by 157 s of intro. Verified: --boot --play --script walks power-on through splashes, ADV, title, (A), main menu, down, (A). A comment above audio.play_bed described the port as CHOOSING the menu track, which HANDOFF Q10 refuted a week ago -- BGM_103 is measured on three independent legs and audio.json says so. Third instance of the drifted-comment trap. The dead phrase is now a check-claims register row, controlled: a planted revival fails and removing it passes. And the boot's wall-clock seconds are a property of this container. ADV takes 146.6 s of wall clock for 137.44 s of media, +6.7%, while S00A runs real time at -0.4%. Not a post-roll and not a general deficit: ADV is 1280x720 and S00A is 768x432, this box has no GPU, and 720p Theora decodes below real time here. The transcode is faithful against a 137.71 s source and the exporter does not rescale. P3/P7 artifacts quote seconds containing that deficit -- reproducible here, not a statement about the port or the game. Comparisons with the Decoder's measurements must go through media length, not wall clock; they carry an explicit emulator pacing factor for the same reason and I had been quoting mine as exact. Their negative result on LOAD GAME, TUTORIAL and OPTIONS leaves guard_focus_scope right to count them UNMEASURED rather than 'resets'. The transferable part is their instrument story: a narrow calibrated reader failed, so they generalised to a whole-frame comparison, which died the moment a crash dialog overlaid the frame while the narrow reader kept working. contract-check is deliberately narrow, individually anchored checks for the same reason, and the temptation after an ANCHOR LOST will be to loosen the matching -- trading a failure I can see for one I cannot. Every asserting check passes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
1ca90bfbd6 |
port: EXTRAS resets, measured -- and being right by luck is not evidence
Ring at 347.5 on entry (MISSION SELECT), 427.5 after one delivery-confirmed DOWN, 347.5 on re-entry with the frame 0.0% different from first entry, screen confirmed by eye because an earlier run was fooled about which screen it was on. Two things settle here. The caveat on extras/initial_focus comes off: MISSION SELECT is a genuine initial focus, because a screen that RESETS cannot have a single-entry reading that is measuring history -- that objection was live only while persistence here was unknown. And focus_persists: false for extras is now written explicitly with kind: measured. Nothing changes at runtime, since the port already defaulted to false; the point is that an absent key and a measured false behave identically and mean opposite things -- 'nobody looked' versus 'the game was watched doing it' -- and only the second is visible to audit-kinds. It does not vindicate how it got there and is not recorded as if it did. For one iteration contract-check ASSERTED extras non-persistence with nothing behind it, the Decoder flagged it, and the measurement then agreed. Their separation is sharper than my own account was: declining to generalise the memory was correct, on the evidence then and on measurement now, since the two screens genuinely disagree -- but encoding 'not measured here' as a positive assertion of the negative was a different move that happened to land. Being right by luck does not retroactively make it evidence. The check is rewritten to rest on the measurement rather than left in place looking vindicated. guard_focus_scope no longer polices 'only main_menu': there is no menu-wide rule to state, since two measured screens disagree. It now states both measured values and counts the screens that say nothing, printing UNMEASURED, not 'resets'. Untested and not built on: OPTIONS, LOAD GAME, TUTORIAL. And nobody can separate 'resets to MISSION SELECT' from 'resets to the top item' -- they coincide, since ptbtn11 is both. The port's value is right under either reading and the reason is not established, which matters the day a screen is authored whose opening item is not its first. 16 kind labels audited clean, 14 controls firing, every asserting check passes. The P5 walk artifact now matches a measurement on both halves rather than one measurement and one default. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
5ff278a5ca |
port: an authored value becomes measured, and a difference-only check gets an origin
The Decoder corrected their own focus delivery: the persistence run's item names were two positions out, from a reader using design-space rows against captures carrying Xenia's chrome and a 1.060 scale. Two things follow. initial_focus_kind moves from authored to measured. NEW GAME on a fresh boot, 2/2 fresh boots, both the first menu entry. The value did not change; its standing did, and the upgrade is not because the measurement agrees with me -- they had said my agreeing with their records was no evidence, which was correct, and this is a direct reading independent of the reasoning that chose NEW GAME here. "First entry" is load-bearing: since the menu remembers its cursor, a reading taken later measures history, which is the objection that voided the earlier TUTORIAL-versus-NEW-GAME disagreement. The superseded reasoning is kept under (was) lines -- the field existing and being labelled honestly is what made arriving at a measurement a label change rather than an archaeology problem, the third time that has paid off after loop_start_s and the +0x08 read. My check_focus_persists anchor survived a correction it should not have been able to detect. It anchors on the heading, the conclusion, not on the item names. That is lucky rather than designed: the conclusion is geometry-free -- ring at y 384.0 before the round trip and 385.5 after, an equality immune to a constant offset -- while the names were not. The check would not have caught the label error, and nothing in it distinguishes anchored-on-a-robust-claim from anchored-above-the- part-that-was-wrong. Their generalisation: a control that only checks differences is blind to the origin. check_splash_dwell is that shape -- it compares the widest gap between keyframe times, and a reader with every time shifted by a constant passes. Added check_splash_times, asserting the absolute list the contract prints. Origin and difference now fail independently. Writing that control reproduced the error one level down: its perturbation literal was written from memory of the prose, with a space where the document has a newline, so it reported its own anchor gone. A control written from a memory of the source rather than from the source is the class of error these checks exist to catch. Thirteen controls, all firing. Q2 closed: fixed same day, and the row was worse than I reported -- the splashes were also mis-paired as 10/11, one half each of two different pairs. EXTRAS remains unmeasured; the run meant to settle it navigated to OPTIONS believing it was EXTRAS. Every asserting check passes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
6f6aea0f5d |
port: audit every kind label, and seven rested on a neighbour's argument
tools/port/audit-kinds reports what each in authored/ rests on. Nothing had ever checked them, which is the point -- the disciplines that fail this way are the ones that never visibly failed. Seven of fifteen labels, every goto_name_kind, had no of their own. Four scored ok on the first run because the audit fell back to the parent's , which argues the DESTINATION while the label is about where the NAME came from. That is the same error I was corrected for the previous iteration, one level down: crediting a claim with evidence that does not bear on it. Borrowed evidence is now its own outcome, and all seven carry a why citing HANDOFF Q4's own words and stating that the port never branches on the field. The audit refuted itself twice first. It counted only paths, shas and filenames as citations, so HANDOFF Q1 and PORT-MISSION section 7 read as citing nothing -- four false positives, and an audit that invents defects is worse than none because its false positives are indistinguishable from its true ones until each is opened. It also resolved paths against committed refs only, failing on a citation to the tool being written. Both fixed. It still cannot read a cited page to confirm it says what the why claims, and prints that every run. MEASURED and measured both existed; a consumer comparing == measured misses the other, and a label that fails to match reads as ABSENT rather than wrong. Normalised. Refutation attempt on HANDOFF Q2's map of GP_TITLE. The headline survives and is exactly right: 4 UI states + 2 loading variants + 2 boot splashes = 8 states shipped twice = the 16 entries the archive holds, confirmed against my export's entry map. But the row enumerates six of those eight -- entries 10, 11, 13 and 14, publisher_logo and developer_logos, appear nowhere in it. A reader counting Q2 gets twelve, and this is the row already corrected once for an ordinal-versus- entry error, which is the mistake four unlisted entries feed. The port is unaffected; both splashes are exported, named and verified at RMSE 2.17 and 3.05. Every asserting check passes, audit-kinds included. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
b3da5c1d48 |
port: correct a check that asserted an absence of measurement as a finding
The pair I shipped this iteration -- focus_persists on for main_menu, off everywhere else -- reported both halves as agreement with the contract. Nothing measured that extras does not persist. The corpus has EXTRAS' opening item from one entry and (B) restoring the PARENT's focus 4/4; neither says what a submenu's own cursor does on re-entry. Caught by the Decoder. It is the mirror of the trap it was written to avoid. I refused to let a derived menu-wide rule overwrite a measured value, then let 'not measured here' become a positive assertion of the negative. Both treat a gap in the corpus as if it carried information and differ only in which direction they fill it. And the failure mode was the bad one: if the game does persist EXTRAS, the check holds the port to the wrong behaviour and passes while doing it. check_focus_persists now asserts only the measured half. The scope became a separate guard with its own outcome word -- 'only main_menu, AUTHORED DEFAULT, unmeasured elsewhere' -- which still fails if widened, since that should be a deliberate edit, but can no longer be read as the game being known to reset. focus_persists_why records the correction rather than being rewritten. It also weakens a label. EXTRAS' initial_focus is marked measured and was taken on a single entry; now that the main menu is known to remember its cursor, a one-entry reading of any screen may be measuring history rather than what the screen opens on -- the same objection that reframed the TUTORIAL/NEW GAME disagreement. The observation stands, its reading as an initial focus does not. Caveat attached, kind left as measured with a note that it changes if EXTRAS turns out to persist. Not building on the non-persistence half until their EXTRAS re-entry run returns. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
6e3347a338 |
port: the main menu remembers its cursor -- a measured P5 defect, fixed and scoped
Measured by the Decoder today: (B) from the menu to the title and (A) back returns to the item you left, not to a default; their control passed first, two delivery-confirmed DOWNs moving the cursor exactly two items before the round trip. The port reset to initial_focus on every entry, so a player who moved to EXTRAS, pressed (B) then (A) landed back on NEW GAME. MenuFlow.enter() now consults opening_focus(), and a new set_focus() writes the memory. set_focus() exists because two call sites set focus -- a cursor move and (B)'s restore -- and a memory updated at only one of them is right until the player uses the other. focus_persists is true on main_menu and nowhere else, and the scope is the authored part. wrap generalised because it was measured on two screens; this was measured on one. Here that is stronger than a preference: extras opens on MISSION SELECT as a MEASURED initial focus, so a menu-wide memory would have silently replaced a measured value with a derived one. Both halves are in one artifact, because a one-sided test passes a port that quietly generalised: the menu returns to ptbtn05 after the round trip, and extras opens on ptbtn11 both times despite being left on ptbtn12. contract-check asserts the pair -- on where measured, off elsewhere -- and fails its known negative. Eleven checks. Not assumed: whether the memory survives a reboot, or whether any other screen has it. Their reach is one boot, one round trip, one direction. The finding also reframes this morning's initial-focus warning without settling it -- if focus persists, a reading not taken on a fresh boot's first entry is measuring history. NEW GAME stays authored, on its own reasoning. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
abaa9de4e3 |
port: check the walk as well as the contract, and a defect I nearly filed off a debug pin
docs/game/navigation.md is a second document unreachable from main, and
authored/flow.json is its executable form -- nothing in the port fails when a
label drifts from it. Three more checks in contract-check, anchored on the walk's
own text: the five main-menu labels in order, EXTRAS' three items, the cursor
wrap. Ten checks now, ten known negatives, all passing.
The manual audit behind them found nothing else: initial focus is already
kind:authored citing Q5's instability, left_right is an explicit no-op,
auto_repeat is measured, unexported destinations are marked blocked with reasons.
Refutation target: the walk's claim that the ring is the ONLY thing moving on the
settled menu. Cannot be tested against the game from here, but can be tested
against my renderer, which is the direction that matters. Five renders across a
full ring cycle: 1428 of 921600 pixels vary, 0.155 %, one 46x44 cluster beside
the focused item. The port animates one ring, not five -- worth checking, since
all five ptbtn01f..05f declare the same 120-unit cycle and a renderer running all
of them would look identical until you diffed frames.
Then I nearly filed a serious P5 defect against myself: sweeping --leaf-time with
the ring pinned moves 10.4 % of the frame, full-screen. It is not a defect. That
pin addresses the build-in -- ptloop01 runs t=0..600, ptloop02 t=0..720 -- and at
settle both park off-screen at x=1521 and x=-839, with loop_leaf_on_screens
scoped to the title alone. The general form: a pin that can address states the
screen never occupies will manufacture defects on demand, which inverts what the
three pins are for.
The +0x08 ask came back answered and is not consumable. ui_layout::loop_length_units
is public at
|
||
|
|
4046c23343 |
port: check the contract's numbers instead of reading 4111 lines of it
HANDOFF on main is 926 lines frozen at 9ca1eb5; the live one is 4111 at
|
||
|
|
6a8b80faaa |
port: the contract I read is 3185 lines shorter than the contract
docs/port/HANDOFF.md on main is 926 lines, last touched |
||
|
|
d7d1fa35db |
port: Q10 correction does not reach me; the register's cost is per-mention
Their stale Q10 row does not touch my tree: stems_why already reads 'a bank is exactly TWO waves of identical duration', the corrected understanding, and the three-sub-waves discrepancy is recorded here as refuted. stems: sum unchanged. Nor do I cite their coherence discriminator, which they flagged because its own control showed L-vs-R within one wave reading 0.22-0.50, so its premise fails in this material. Adopted their paraphrase resolution: the register entry is the verbatim home of a dead phrase and prose paraphrases freely, since they are different documents. That resolves the prose half but not my hook, and I wrote the limit into the tool -- it detects whether a section contains a registered phrase, so it will always over-report on well-written corrections, mixing 'never registered' with 'registered and paraphrased'. A prompt to check, never a defect count. Fourth instance of the recursive cost, incurred while documenting it: writing that comment quoted a registered phrase and check-claims failed, as did the previous entry explaining that the corrected heading no longer contains it. Both marked. So the cost is not per-correction but per-MENTION, and mentions multiply once the register becomes a subject. Four instances, each inside text about the mechanism. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
653e317197 |
port: full regression passes; the phase term moved two published rows
Ran the suite after a session of edits to boot.gd, screen_view.gd, four tools and two authored files. Every asserting check passes, and verify-screen's two DIFFERS are the named pair with per-screen reasons. Two oracle rows moved: title_plate 12.83/0.00% to 13.04/0.09%, title_band 15.31/0.35% to 12.86/0.00%. Opposite directions, which is a phase change rather than a regression, and the cause is mine -- adding --leaf-time=0 to verify-capture's render sites pinned the sweeps while the captures froze them wherever the shutter caught them. That makes the capture-phase term concrete: I documented +/-5.56 for title from a sweep, and here it moved two published rows from a one-line harness change. It also touches a number I published -- the boot-end-frame 0.00% was measured before the pin, and the equivalent row now reads 0.09%. Both inside the term, and the right reading is that neither is 'the' number. Also narrowed the withdrawal-time hook. Its regex matched headings ABOUT corrections rather than headings making them, so 33 was a measurement of the regex; narrowed to a leading WITHDRAWN/CORRECTION/Refuted, it gives 10, all genuine retractions. Residual limit named: several of the ten are flagged because the registered phrase does not appear in that section -- the corrected JP heading reads 'does NOT go against the port', which does not contain 'goes against the port'. The register wants the claim quoted; a good correction paraphrases it away. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
8baa303d34 |
port: build the withdrawal-time hook, and violate the rule it enforces while writing it
They ended with 'it needs a hook at withdrawal time, not a sweep'. Expressible, because a correction here has a shape: a heading carrying WITHDRAWN / CORRECTION / refuted. A correction section containing no registered phrase is a death argued and never indexed. check-claims now reports them, and the first run names more than my 'four of eight' -- the shortfall runs back through earlier work. Reported, not asserted, deliberately: not every correction retires a claim, and forcing rows for those would push rows in to silence the check. Two failures while building it. The first version pasted the register rows into its own heredoc, so every registered phrase became an unmarked quotation and check-claims flagged its own source -- a tool violating the rule it enforces by being written. Fixed by passing the register through the environment. And writing up the previous catch re-introduced three unmarked quotations: describing a refuted claim quotes it, so every correction is a new occurrence needing the token. The cost is recursive, which the header implies but does not say out loud. What the hook does not do: it fires when a correction is written, so it closes the gap between arguing and indexing, not between believing and arguing. Nothing here would have caught me copying their 'structural' claim into my record. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
a0f27ec69b |
port: their REFUTED gap, in a register I had and fed nothing
Their finding: eight claims died this session and none reached REFUTED.md, the file their brief says to grep before proposing anything. The pages are where a refutation is argued; the index is where it is found. Mine is the same gap and worse in one respect. tools/port/check-claims is a register that FAILS the run if a refuted claim is quoted without its [refuted] token, and it is in check-all -- so an entry enforces rather than merely publishes. It held 7 rows, all from earlier work, and I added none while withdrawing about 8 claims this session. Registered four. The checker immediately flagged three still asserted unmarked, and every one was inside a correction I had written myself -- the headings-audit table rows explaining the withdrawals, and the EXTRAS withdrawal block. That is the token doing what phrasing cannot: all three read as corrections to a human and the marker fired anyway, because it tests for a token an author places rather than for language that sounds retracted. Marked; the register now passes. Scope: four of roughly eight registered. Not registered -- the compactness precondition, the half-rate defect, 'the eras render identically', and my 16/16/18 rule -- each argued in its own correction and findable by nobody. Stopped at four because each row costs marking every existing quotation by hand. And nothing mechanically checks that a future withdrawal reaches the register. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
a84eb14900 |
port: close their XPR lead, and find their class in the lane I called clean
They flagged five XPR_* texture toggles as relevant since I consume textures, and my off-edge splash residual -- non-tonal, ~0.5 RMSE above quantisation, no candidate -- has the shape a subtle decode difference would produce. Closed: the toggles live in texture.rs::decode_surface, shared by from_xpr2 and cube_faces_from_xpr2, and my exporter calls neither -- sprites come from t8ad::parse. t8ad.rs reads no environment variables in its 202 lines, so the sprite path has no hidden freedom either. The candidate is eliminated with no replacement. Enumerating what my exporter reaches turned up SYLPHEED_KF_TIME_SHIFT, which they reported as absent from crates/. True on their branch, false on mine: my ui_layout.rs is the stale era and the knob is live at line 497. The pinned tag has 0 occurrences (2 of LEGACY) so export/ cannot be perturbed, but verify-screen builds its reference from the workspace, which can. Tested both directions: with the knob the reference reports rest t=12, the corrected reading, and the era guard passes; without it, t=70 and the guard refuses. So the knob is the working remedy that makes a workspace-built reference usable, and it appeared in no tool, help text or instruction in my tree -- their exact class, in the lane I had just told them was clean. The refusal message now carries the remedy. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
66f0adce02 |
port: sweep instructions above descriptions -- the silent class is clean, two loud hits
Their sharpening: a stale instruction manufactures a false confirmation, strictly worse than a stale description that merely misleads. Applied to my instruction surface, the documented invocations in tool and script headers. All fifteen distinct flags across those examples are parsed, so nothing in my headers can produce their failure mode by being inert. But 'parsed' is a proxy and its gap is known -- --shots parses and does nothing on the --boot path -- so I ran two documented examples end to end rather than trusting the grep, and both produce a 1280x720 frame. Two hits, both loud rather than silent: 11 references to tools/verify-capture and tools/verify-screen, paths that do not exist since the tools are under tools/port/ (fixed in 4 files); and check-all claiming eleven tools where there are fourteen (now states both so the sentence dates itself). The distinction worth recording: mine fail loudly, theirs failed silently. A wrong path announces itself; an inert environment variable returns a clean wrong result. Both are stale instructions and only one manufactures evidence. Honest limit: I tested the flag surface plus two examples end to end, not all thirteen documented invocations -- the --boot ones take 156 s each. That is a judgement about cost, not a claim of coverage. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
d725f8e2f8 |
port: the dead-rule grep found two more, and the cause is my correction habit
Their generalisation of my 'untimed' marker -- search for the vocabulary the dead rule needed -- is the cheap version and it works. Swept for the nouns of every rule refuted this session. Two real hits: verify-screen:57 still asserting 'all four are COMPOSITED rather than standalone', the reading withdrawn after they tested it disc-wide at 7.9%; and boot.gd:197 opening with the pre-fix 'no time slot' claim before retracting it. Third and fourth instance after spin_period_units and exit_ramp_units, and in all four the correction sits below the false claim in the same block, with both written by me. The diagnosis is a habit: my corrections are ADDITIVE. I append a CORRECTION block and leave the original standing, which is right for a record and wrong for a statement -- a reader takes the first assertion and the retraction three lines later has already lost. The habit that creates these is the same one I adopted to make corrections honest. Fix: keep quoting the original but demote it grammatically, leading with 'what this used to say'. Both rewritten. Verified comment-only by artifact rather than by reading -- the main_menu render is byte-identical before and after. Also records agreement with their caution: the failed gap+clear rule was rejected, not narrowed to menu transitions. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
02ab62e28d |
port: sweep my own tool headers after theirs -- two hits, both in verify-dwell
Their audit found one defect in sixteen commands and their point that doing one and stopping is the failure applies to me: I had fixed verify-screen and verify-capture and gone no further. Hit 1: verify-dwell built its target as oracle span + the GAME's black gap and scored the port against it, correct only while the port inserted that gap. It does not -- black_hold_units went to 0. On publisher_logo the port runs 0.131 s below the unslacked target, absorbed into an 'agrees' by 0.15 s of slack that is larger than the omission it hides. Hold now read from authored/timing.json; the game's gap printed as its own term. Hit 2: the tool carried '4 presented frames at 2.284 units/frame'. The number is right but it is the disc used as its own clock on ONE capture that ran at 13.1 fps against ~28 elsewhere. Stated bare it reads as a general rate and would contradict Q1's 2 units per rendered frame, a different quantity at normal speed. The derivation was in DECISIONS.md; the tool inherited the value alone -- exactly their defect, and their 'print the population beside the number' fix applies unmodified. Not found elsewhere: check-capture's percentages all name their population; check-claims, check-modding, index-decisions and strip-padding assert no measured quantities. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
edad692500 |
port: verify-dwell built its target from the GAME's black gap while the port's is 0
Audited my own tools the way they audited theirs. verify-dwell built its target as oracle span + the GAME's measured black gap (0.114-0.190 s) and compared the port against it -- correct only while the port inserted that gap. It does not: black_hold_units went to 0 three iterations ago. So the port is expected to run short by the gap, and on publisher_logo it does -- 0.131 s below the unslacked target, which the 0.15 s wall-clock slack was quietly absorbing into an 'agrees'. A verdict that passes because the slack happens to exceed a known omission is not a verdict. The hold is now read from authored/timing.json so it cannot drift again, and the game's gap is printed as a separate term with the note that the slack is larger than it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
d8488640a8 |
port: my backdrop predicate is exact in GP_TITLE and its reading was wrong
I offered 'a declared opaque-black backdrop distinguishes standalone from composited' and asked for it to be tested against archives I do not have. It was. The split reproduces exactly: derived independently from the disc, GP_TITLE gives 12 with and 4 without, the four being entries 0-3 -- my build_00, build_01, press_start, press_start_jp -- with element names matching. Two genuinely different paths, my export against their disc reader. The reading does not survive. Disc-wide the predicate is rare, 76 of 965 builds at 7.9%, with GP_HANGAR_ARSENAL 0 of 390, GP_OPTIONS 0/14, GP_PAUSE_MENU 0/6. Read as 'composited' it makes 92% of the game composited, which the archives do not support. What survives is narrower: it separates screens that BEGIN FROM BLACK from everything else, and their sharpening is the part I would not have reached -- the negative class is heterogeneous, so a two-way rule cannot express it. My caveat named the exact test that refuted the reading, but I still put the refuted interpretation into verify-screen's header as a stated fact while the hedge lived in DECISIONS.md. Corrected, with the 7.9% figure and an explicit do not carry this into the four unexported archives. Hedging in the write-up does not protect the claim shipped in the tool -- the same delivery gap as the capture-phase term, repeated four iterations after fixing it once. Within GP_TITLE the rule is exact and --black for those twelve is justified from the file rather than assumed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
c6735f55a6 |
port: audit the --black premise -- declared on 12 screens, assumed on 4, all composited
Their finding that screen render --black's premise is declared on the splash builds is checkable across my whole export, and verify-screen passes --black to all sixteen screens on that premise. Audited by asking whether a screen declares a full-screen untextured primitive at t=0 with fade_argb 0xff000000. Twelve do -- pteff00 on both titles, both menus and both extras, palogo_eff0 on all four splashes, pgloading_eff00 on build_12/15. Four do not: press_start, press_start_jp, build_00, build_01. All four exceptions are composited rather than standalone. press_start is one element, the plate, whose own name_why records it is composited over the title. build_00/build_01 carry the pgloading_* set without the pgloading_eff00 backdrop that build_12/15 declare. Harmless where used: verify-screen gives --black to both renderers so the assumption cancels in a consistency check, and verify-capture already scores the plate over the title rather than on black. The exposure was real and the tooling had already routed around it, which could only be established by looking. The rule that falls out: a declared opaque-black backdrop distinguishes a standalone screen from a composited one, derivable from the file rather than from a name. Recorded as a rule with its evidence -- sufficient as observed, not proven necessary, on four exceptions. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
51c85ed691 |
port: print the capture-phase term beside the numbers it qualifies
Their closing point -- the thread lived in messages and docs/re/, which by our own rule means it was not delivered -- applies to my side. The capture-phase term was in DECISIONS.md, but verify-capture is what prints the numbers it qualifies and it said nothing: a reader saw title 14.16 with no sign that +/-5.56 is inherited from where the shutter fell. Now printed per row: title +/-5.56 regression only, main_menu +/-3.78, extras +/-3.73, and both splashes marked as carrying no free-running element and meaning what they say. Header records that --leaf-time=0 is a convention, not the game's phase. Also names a gap their own update exposes: they landed the leaf facts in HANDOFF, correctly, but HANDOFF as I read it contains none of them -- their work is on auto/build-ordinal-audit and origin/main is 145 commits behind. So the facts reach me only through messages, the channel the rule says does not count. Writing it in the contract is necessary and not sufficient when the contract lives on an unmerged branch. My BLOCKED.md and DECISIONS.md carry the status sourced to their sha so my tree does not depend on a HANDOFF I cannot see. Second structural consequence of main being stale, after the Cargo.toml pin. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
49958ff090 |
port: the third clock was in my own enumeration and I did not wire it
Last iteration I enumerated three free-running clocks, wrote that the leaf is pinned only by --leaf-time, then tested reproducibility without passing --leaf-time and concluded nothing free-runs on the menu path. The answer was one paragraph above the experiment that contradicted it. My own flagged weakness found it: deliberate wall-clock variation via --script=wait:N, putting the capture at t=96 units against t=369. Spin pinned only, wait 0.5 vs 5.0 differs by max 91.19 per channel; with --leaf-time=0 added it is byte-identical. draw_leaf_for is ptloop01/ptloop02, present on main_menu and not just the title, which is why that row drifted. verify-capture passed --loop-phase=0 and not --leaf-time=0 -- I fixed the clock I had been bitten by and left the one I had merely listed. Enumeration without follow-through fails exactly like no enumeration. Both are now pinned at all six render sites. main_menu returns 13.21 across three runs and two renders after different waits are byte-identical. The number moved 13.26 -> 13.21 and that is NOT an accuracy improvement: pinning the leaf at phase 0 puts ptloop01/02 at one specific pose rather than wherever the wall clock left them. A different configuration, now reproducible. Which pose the game shows at rest is not settled by this. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
cf8f001956 |
port: the oracle harness was nondeterministic and I quoted it for a dozen iterations
verify-capture's main_menu row reads 13.30 / 13.27 / 13.25 / 13.26 across runs this session while every other row is identical to the digit. I cited those numbers repeatedly, including in the rest() adjudication. Cause: the focus ring spins on time_units raw rather than the pose clamped by holding -- deliberate and correct, since the ring is the one thing on a settled screen that keeps moving -- so its angle at capture is set by the wall clock. extras is stable because nothing there spins. --loop-phase already existed and did not cover it: it pins the looping focus record phase, while the spin is a second free-running clock I guarded once and never connected. Extended loop_phase_units to pin the spin too, and verify-capture now passes --loop-phase=0 at all four render sites. The control matters because the drift was intermittent -- three unpinned runs gave 13.25, 13.26, 13.26, so three pinned runs agreeing would prove nothing. Phases 0/30/60/90 give 13.2583 / 13.1991 / 13.2637 / 13.2588: the pin is live and the 0.065 spread is the whole of the observed drift. Non-finding recorded so nobody mines it: phase 30 scoring lowest is not evidence about the ring's real phase -- 0.065 against a ~13.2 gamma floor is 200x too small. A margin only means something against the noise it sits on. No conclusion changes: the smallest margin any of them turned on was 0.14% differing area. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
dbbf28e22f |
port: my branch IS the stale era, and verify-screen's reference was never its own build
Told the Decoder their diagnosis was wrong. They were right. ui_layout.rs is md5 b6c19d08 in my working tree, at HEAD, on my pushed branch and on origin/main -- one file, stale marker present, tree clean. What misled me is the same trap a third time: CARGO_TARGET_DIR is a shared /sylph-home/port/target-container, so two source trees write one binary and cargo fingerprints per source path -- each build reports Finished while the binary on disk belongs to whichever tree wrote last. A CLI built from my workspace is 3a39fce (stale, rest t=70), identical to one built from origin/main; the binary verify-screen actually used was 8e0aa76 (fixed, rest t=12), from a tree nobody had named. It happened to be the right era, which is worse than wrong -- it agreed with the pin by luck and one rebuild would have flipped it silently, and title_jp differs by 74507 px between eras. verify-screen now reads the reference CLI's pteff00 rest instant and compares it against the export the port reads, refusing to score if they disagree. Controlled both ways: passes with the matching binary, refuses the stale one built from my own workspace. And the pin is load-bearing, not an annoyance to revert: the workspace crate is stale, so the pin is the only reason the export is correct. Consequence worth stating -- my published branch carries the stale crate, so anyone building sylpheed-cli from it gets the stale decoder. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
8666c33a6a |
port: WITHDRAW the 'eras render identically' measurement -- I compared a binary with itself
Last iteration I overturned check-all's allowance on a measurement of 0 pixels between the two decoder eras, and rewrote the tool's reason around it. The two binaries had the same md5: one built in a worktree at formats-pin-2026-08-30 and one from the workspace, and both commits carry the record-layout fix. I compared a binary with itself and reported the zero as evidence. The 508-line diff I cited was real and irrelevant -- it does not straddle the fix. Done properly against origin/main, verified stale by the Decoder's own control (rest t=70 vs rest t=12) and by differing md5s: title 0 px, main_menu 0 px, title_jp 74507 px -- reproducing their figure exactly, under their flags and mine. My second hypothesis, that --animated masked it, was also wrong. What survives: the era still cannot explain this script's rows, for a fact I had not established -- both sides of the comparison are the FIXED era, since a binary built from the pin and one from the workspace have the same md5. Right answer, wrong evidence. The note now carries its condition: title_jp is era-sensitive, so if the reference is ever built from a different era than the pin, that row's cause changes. Twice now a correct conclusion has come through a broken experiment, and both times the tell was two things that should differ producing identical output. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
ecd5e56e0c |
port: check-all excused two failing rows with a measurably false reason
The suite reported '2 DIFFERS, allowed: the pin is not on main, so this compares two decoder eras', and I had quoted that for several iterations without testing it. Built sylpheed-cli at formats-pin-2026-08-30 and at workspace HEAD and rendered through both: title, title_jp and main_menu come out 0 pixels different, despite 508 lines of difference in ui_layout.rs. The eras are not the cause, and the allowance was excusing a real signal with a wrong explanation. A second defect in the same eight lines: the expiry tested formats-pin-2026-08-29d while Cargo.toml pins formats-pin-2026-08-30, so it would have expired on a tag this tree does not use. The real reasons are per-screen and already documented: title is the ptloop sweep phase residual, title_jp is the --pose=rest sparkle handling -- where the port's shipped pose scores +0.9994 against the game to the reference's +0.8727, so the port is closer to the game on the row the script calls a disagreement. Replaced with a named set: title and title_jp by name, any other DIFFERS fails. A count cannot notice a different screen drifting while the total stays at two. Controlled both directions -- passes on the known pair, fails on main_menu or extras. The pin reminder now reads the tag out of Cargo.toml so it cannot drift. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
d613609aaf |
port: check-all hung for an hour on an ffmpeg that had already finished its work
check-all sat on two lines of output for over an hour. The cause was the 5.1 bed in check-capture-controls: ffmpeg completes the filter graph and then never exits. Diagnosed rather than guessed -- the output reaches 4604262 bytes, exactly 8.0 s of 5.1ch/16-bit/48kHz, the full intended length, with the artifact correct on disk while the process hangs. Three formulations all hang and all produce byte-identical output: the original, one with -t 8 bounding the output, and one with explicit asplit feeding each atrim (the textbook fix for multi-use of a single input). So it is not the split, not the output stage, and the artifact is not in doubt. Worse than the hang: it leaks. An orphaned ffmpeg from this script's earlier aloop form was still running after 9.5 hours, burning CPU across runs nobody was watching. boot.gd's header already names the shape -- a job that waits forever reads as a job still working. Bounded with timeout, and the ARTIFACT is now checked rather than the exit code: the bed's duration must be 8 s or the sweep refuses to score itself. That is the better test regardless of the hang -- an exit code says ffmpeg thought it was done, the file says what it wrote. The step now completes in 99 s and the sweep matches its specification. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
a43dee3ab0 |
port: the loading screens are no longer black -- it was the paint order
verify-screen's header has said since P1 that build_12/build_15 render pure black in both renderers, with an open question whether that was the port's bug or the decoders' reading of rest. Measured today: max 214.5 on both sides, mean 1.949 port against 1.918 reference. Not blank, and they agree. It was the paint order. My own earlier measurement had already answered it and I had not connected them: removing the forced-backdrop pass makes the first element pgloading_loop5 and the black screen returns. pgloading_eff00 carries layer: null, layer_source: none -- the only elements in the export with neither a read nor an implied key -- so its position rests entirely on the occlusion constraint. The guard stays, with the stale paragraph kept as history. It was right when written, and a guard that stops firing is the kind that rots out of a tool. Refutation attempt on the Decoder's census scope: my six transient ptlogo_back2eff* on title are also GP_TITLE, so if they were fallback fires their count of four would be wrong. Their claim survives -- all six reach rest by the plateau path, alpha 255->255 with identical pos and scale, so the fallback never runs. The two censuses differ in scope, not in fact. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
835acf930e |
port: WITHDRAW the claim that the port drifted away from the game -- wrong frame
The previous entry scored verify-screen's title_jp frame against the oracle and concluded the port had moved away from the game. That frame is posed --pose=rest, which the port does not ship. Posed as it runs, the disputed block scores +0.9994 against the reference's +0.8727, and the whole surface +0.9652 against +0.9200 -- holding under gamma compensation and on the English control (+0.9946 vs +0.9560). The port is closer to the game than the reference on both title screens. Mechanism: ptlogo_back2eff1 is (0,0)(98,0)(100,255)(102,255)(104,0) -- a 4-unit sparkle whose rest.t is the peak of its own flash. Six of them stagger across the logo, so --pose=rest fires every sparkle at once. The 25.6% excess light was real and was in a frame nobody sees. verify-screen is not at fault: it poses rest deliberately, so that both renderers read one decoder and the run is a consistency check. I used a consistency-check frame for a correctness question. Its header now says its frames must never be scored against a capture. A second claim in that entry was also wrong -- both screens draw those layers under pose=rest; I had compared a --menu timeline log against a verify-screen rest log and read a mode difference as a screen difference. verify-capture takes a fifth per-row field, a capture crop, because this capture is a full display frame with the surface at +0+45 while the others are pre-cropped. With it title_jp reads RMSE 20.91 / 1.04%, beside title's 14.16 / 0.21%. The row prints 'no capture' until their branch merges. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
8ae0ec2287 |
port: verify-screen was nondeterministic; pin the pulse phase in the harness
Running the full set after the plate fix, press_start returned over3 5021, 8919, 5021 on three identical runs. The plate's looping focus record takes its phase from time_units, which free-runs, so the captured frame lands wherever the grab fell -- while the reference renderer cannot pulse at all. The port is not the thing that is wrong: the pulse is measured and a thing that pulses does not stop because the screen arrived. ScreenView.loop_phase_units pins it, negative means free-running and stays the default everywhere, and only the harness passes --loop-phase=0. Controlled: pinned, 3 runs identical; free-running, 3 of 4 identical and one different. That 3-of-4 is why it survived -- it looks deterministic most of the time, and without the negative control a no-op flag would have been indistinguishable from a fix. With the phase pinned press_start reads max 1 / over3 0 OK -- the recorded baseline exactly. Fifteen of sixteen rows now match. The sixteenth, title_jp, has genuinely drifted: 155/20498 -> 233/61208, deterministic, on the Godot side, localized to one 350x396 block at (405,74). There is no capture of the Japanese title, so I can say the renderers moved apart but not which moved. Recorded as an ask, not resolved. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
71243bcc60 |
port: confirm no screen of mine carries a .tbm, and guard verify-screen anyway
The Decoder found that sylpheed-cli screen render silently omits the background of any .tbm-bearing build, and stated that none of my screens has one. That is a claim about my tree and it decides whether my regression baseline is sound, so I tested it: zero .tbm across all 16 builds in my manifest -- wider than the five they said. Both controls fired (GP_TUTORIAL build 0 -> pubase.tbm; GP_TITLE build 5 -> none); my first attempt's control printed nothing and I nearly read that as agreement. verify-screen now names the omission on any .tbm-bearing row. It cannot fire on a screen I ship -- which is how a guard goes dead -- so its expression is controlled directly in both directions. No verdict or bar changes. Regression unchanged: title max 6 / over3 790, main_menu max 4 / over3 0. Their identification (reading TUTORIAL off the framebuffer) and my edge correlation (run before their message, blind to the text) agree on GP_TUTORIAL build 0 from no shared assumption. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
8da453478d |
port: a refuted-claim register, enforced by check-all
The Decoder's audit of their own corpus found four refuted claims standing -- including one they had corrected to me, agreed with, and written a METHOD entry about, without landing it for a full iteration. A hand audit finds what is there on the day it runs; it does not stop the next one. check-claims is a register: every occurrence of a refuted claim must carry an explicit [refuted] sentinel within 400 characters. It found four more unmarked occurrences than my manual pass had, including one in authored/audio.json. The marker is a sentinel rather than a keyword because the first version's every failure was a quotation inside a correction whose wording lacked the keyword. The temptation was to widen the window until they passed -- tuning a threshold until the answer comes out right, in the tool built to catch that. 21 quotations marked by hand; proved it fails by removing one. Also fixes the Decoder's other finding in my corpus: BLOCKED's voice row had a struck heading with three sentences below still asserting in the present tense. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
81319ea20e |
port: the dead-press check was passing by luck -- diagnosed and fixed
Two iterations ago verify-menu-audio's bit-identity assertion began failing and I filed three suspects in the port. It is none of them. Three IDENTICAL invocations give two outcomes, 1.207438 s and 1.300317 s, differing by exactly 4096 samples -- one mixing buffer. The recording quantises to whole buffers and a one-buffer shift moves the length and alignment of everything in it. The premise -- cross-run bit-determinism -- was never guaranteed. It held while timing sat away from a buffer boundary, and a larger export moved it onto one. A test that passes by luck reports the luck running out as a regression in the code, which is what it did: two iterations of suspects, and the port was never involved. The fix keeps exact equality and no threshold, allowing the comparison to slide by whole buffers -- the one degree of freedom the recorder has. Proved it can still fail: ctrl against walk differs at every alignment. Distinct from the earlier entries: this check ran and answered the right question, resting on a property of the environment nothing verified. State what an assertion assumes about the machine, not only what it checks. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
8aea939050 |
port: audit for findings living only in code comments; found the mirror trap instead
The Decoder lost a finding whose only record was a script comment and asked whether I have the same. Audited every measurement-shaped token in comments across the exporter, the GDScript and the tools against everything in docs/. Seven candidates, six were my matcher (thousands separators, ranges written differently, precision). The findings are all in DECISIONS, including the leaf comment's capture-measured centres and the 11.5 px residual. The one real defect is the opposite: check-capture's control table and AUDIO-VERIFICATION.md had DRIFTED -- 53.3% against 53.2%, twice each, for one control whose file is gone so neither can be re-measured. They lost a finding to having one record; I lost a digit to having two with nothing keeping them equal. Fixed by citing rather than restating. Also corrects a message: I told them my computation reproduces their published centres to half a pixel. True, and MODEL against MODEL -- against the capture this corpus already records 992.0/467.2, an 11.5 px residual. The half-pixel agreement is two derivations of one model, the correlated-instrument shape I have been careful about all week and did not apply to my own message. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
477cc6e6a7 |
port: index DECISIONS.md -- it already answered last iteration's question
Last iteration I filed title and title_jp's disagreement with sylpheed-cli as mechanism-unknown, to the Decoder as well as here. Both were already explained in this file, under headings that name the two screens. Checked rather than assumed. title: still ties on 0x8083, 0x80a0 and 0x8010, and the export declares paint_order_ties unresolved; the old entry's 904 px in the glow band matches my 790 px at the same place, same 4-6/255 magnitude. title_jp: the 'only non-integer scale' claim finds 26 keyframes export-wide, but exactly ONE element visible at rest -- ptlogo_eff2 at 125% -- which is the pose verify-screen uses. It survives narrowly. The failure is navigability: 6502 lines, 111 sections, no index, so 'has this been decided?' had no cheap answer and re-deriving it looked like diligence. index-decisions generates the contents; check-all runs --check. No line numbers (the first version was a fixpoint that failed its own check, and appends would invalidate them all), and checked, because a stale index answers 'already decided?' with a confident no. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
fd6d9b6d47 |
port: add check-all; verify-screen ignored its own statistic; 'six expected DIFFERS' was wrong
Eleven tools and nothing ran them together -- the ninth instance of correct, documented and unexercised, one level up. check-all runs the four that assert, reports the oracle table, and gives verify-screen an allowance that EXPIRES when the pin lands rather than standing forever. All eleven exercised first; none had rotted. verify-screen computed over3 because 'a single max cannot tell 2 pixels from 25 444' and then decided the verdict on max alone: main_menu (max 4, over3 0) read DIFFERS while extras (max 3, over3 0) read OK. The bar is unchanged; a frame with no pixel over it now gets its own ROUNDING verdict. And corrects a claim I have given the Decoder more than once. The real count was ten, now eight: six forced-backdrop, two rounding, and TWO UNEXPLAINED -- title at 790 px and title_jp at 20498, neither carrying a forced element. My leaf hypothesis is refuted: emptying draw_leaf_for changes the numbers not at all. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
8b4ec4414d |
port: make the documented control sweep executable; it was prose
AUDIO-VERIFICATION.md calls its six-file sweep 'the tool's real specification' and nothing ran it -- in a tool whose own history is two invented thresholds caught only by controls. The same document states the principle it was breaking: a control that does not execute is not a control. tools/port/check-capture-controls rebuilds five of the six and asserts their verdicts. The starved capture is gone and is reported MISSING rather than omitted, and deliberately not synthesised from its published statistics -- a control fitted to the answer it must give is not a control. Two things the sweep had to learn to be honest about. check-capture emits TWO verdicts and the doc's table compresses them; the voice control is PASS on channels and UNJUDGED on starvation by design, so the sweep asserts the pair. And a starved file short-circuits before the channel check, recorded as n/a rather than FAIL -- the check did not run and the check failed are different facts. My first 'real music bed' control was -ac 6 from a stereo source and FAILED correctly: an upmix leaves channels silent and byte-identical, which is what the provenance check exists to catch. The control was wrong, not the tool. Rebuilt from six non-overlapping spans of real audio. A second attempt used aloop=-1 and hung ffmpeg indefinitely. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
606eee8f23 |
port: check the five MODDING rules, and label the generated files in the asset tree
MODDING.md calls modding a constraint on the exporter TODAY and nothing verified it -- the same shape as the black hold, skipped[], stop_bed and --focus. All five rules pass, so check-modding is a guard rather than a fix, and it is proved able to fail: a stripped .cmd header, a bogus.bmp, and one orphaned PNG each exit 1. It found one thing: the .cmd encode-cache sidecars sat in the modder-facing tree with nothing saying what they were. They now carry a header. The header is excluded from the cache key so rewording it does not re-encode four minutes of video, and the sidecar is refreshed whenever its text differs rather than only on re-encode -- otherwise a header change could never reach an existing export. Also partly answers my own question to the Decoder: there is no general capture-path floor, because the port matches live-title-press-a at 0.00093% full-frame and 0.000% across the band. The 0.301% is specific to that pair. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
e811bb99e2 |
port: place the last unused capture; its residual is oracle-to-oracle, not the port's
live-attract-title-press-a-band.png is 1279x120 and the harness could not compare a band. Placed by sliding: y=520, a 25x drop over five pixels, and it fits at t=236-238, the plate's own window. Its 0.354% is not the port's error. The port reproduces the same band of live-title-press-a EXACTLY (0.000%), and the two captures differ from each other by 0.301% -- two thin strips, 248x5 and 206x1, the shape of a sub-pixel edge difference. The row's job is to stay near the oracle-to-oracle gap, not reach zero, and it says so. I had begun writing that the attract-returned title differs from the boot title. It is two hairlines. The connected-component breakdown stopped it. All eight live captures are now used. The three that were idle were each blocked by the harness, not the capture. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
a758f9b247 |
port: --focus was ignored on the menu path; focus rendering now verified against the oracle
live-main-menu-options-focused.png -- the only capture of a known focus state -- was untestable because --focus= parsed, was stored, and was overwritten by the authored initial focus on every _menu_enter. Every run logged focus ptbtn01 whatever was asked for. Now pushed into the menu model so navigation continues from where it was forced. With it working, each capture picks out exactly one button: ptbtn04 at 0.1355% against 0.70-0.82% for the others on the OPTIONS capture, and ptbtn01 at 0.0705% against 0.72-0.84% on the plain one. 5x and 10x discrimination. First time the port's focus rendering has been checked against the game at all -- the existing main_menu row uses an authored focus and could never have caught a focus error. Records in flow.json that live-main-menu.png shows NEW GAME focused, so the authored initial_focus matches the one frame it can be checked against -- and that this does NOT overturn Q5's measured instability. It stays authored. Adds main_menu_options to verify-capture at 0.13%. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
69e9043b96 |
port: a second capture closes the sweep-geometry question; title+plate matches at 0.00093%
live-title-press-a.png was unused in the corpus. Posed at t=237 -- inside the plate's 8-unit window -- the port matches it at 0.00093%, against 0.0124% for the no-plate capture at leaf phase ~400. Two captures, two different phases, both under 0.013%: a systematic sweep-geometry error would leave a floor in both, so last iteration's caveat is closed. Sweeping the whole screen's instant against capture 1 gives at best 0.148% at t=230 -- 10x worse than the leaf-only fit. So that capture is the screen SETTLED with the sweeps still looping, which is the first independent evidence for the authored loop_leaf decision. Fixes the cause of a flat 1% floor: --screen=X --overlay=Y pushed the raw elapsed clock into the overlay (9 units at capture), so press_start drew nothing -- the flag whose purpose is 'put the plate on the title'. A static overlay now poses at its own arrival; the --boot shared clock is untouched. Adds title_plate to verify-capture at 0.00%, the most sensitive row in it. Its instant is FITTED and labelled as such. Also records that I nearly committed a wrong cause for the overlay bug: I wrote that nothing drives the overlay's clock outside a sequence. It is driven, every frame, from view.time_units. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
96fe0eac51 |
port: add --leaf-time, and the title's residual is the sweep phase (~400 units, not 357.7)
The Decoder's refined sweep fit had never been testable: verify-capture passed it as a whole-screen --time that pose_at discarded, and asking for it honestly poses past the title's group end. --leaf-time separates the leaf's clock from the screen's. Controls: the renderer is deterministic (3 runs bit-identical) and the sweeps move 0.40% of the frame between phases, so the comparison can see them. Sweeping the full 600-unit span gives a sharp basin at 390-415 units (0.0124%) against 0.2532% at t=357.7 -- 20x. So the title's 0.21% residual is the sweep phase, not structure: at the fitted phase it matches the capture as well as the splashes do. NOT adopted: the port loops the leaf freely and re-posing the harness to the fitted value would be tuning until they match. Filed instead, with the question of whether 357.7 and this are even the same quantity. Also verified last iteration's settle-window change was surgical: only press_start and its twin moved, 14 screens unchanged including title's Decoder-confirmed [160,236]. Settle-window ties exist on 4 screens but all sit under the 30-unit bar, so the arbitrary tie-break never reaches the runtime. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
3c71962698 |
port: the PRESS (A) plate could not be drawn at any instant -- four faults, and a misquoted number
1. --time= was silently ignored on any screen with a settle window >= 30 units: pose_at overwrote the requested instant with settle_instant. ScreenView.frozen now marks an explicit instant and skips both clamps. 2. press_start's settle window was [0,214] -- the dead stretch BEFORE the plate exists -- so its settle instant was t=107, where the element is alpha 0. The exporter now rejects intervals in which nothing is visible. title keeps [160,236], the interval the Decoder's draw stream confirmed. 3. My authored looping_focus_records entry for press_start/ptbtn00 drew a dim focus record INSTEAD of the plate's own sprite: max 0 vs max 252.5. Deleted -- an authored guess that overrides a decode with a worse answer is removed. 4. verify-capture passed --time=5.9617 for the title and it was never applied. Every title figure it has printed, including the 0.26% quoted to the Decoder, was measured at the settle instant under a note claiming t=357.7. Both rows now pose by omission and the note matches. title is 0.21% honestly; splashes unchanged at 0.01%. The boot's end artifact now contains the plate (region mean 95.7 vs 33.6). Corrects last iteration's BLOCKED row, which had the entry's effect backwards. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
7ffcf743a2 |
port: P6 gate verified with sound on the bus; tighten the backdrop guard to a positive primitive test
verify-menu-audio records the Master bus over the P5 walk under the Dummy driver. A dead press is bit-identical to the bed alone; all three cues match their exported wave in the recording with margin over a bed-only control; the cue order matches the script order, which the correlator was never told. The first version of this tool counted envelope bursts above a multiple of the bed and gave 4 cues on one run and 0 on the next from the same script. Replaced with template matching, which has no tuned constant. Cue LENGTH is deliberately not asserted -- the bed masks the tail and I nearly filed that as a defect. Also acts on the Decoder's .tbm self-refutation. No port verdict is affected -- all six forced elements are .prm solid black, and GP_TITLE has no full-screen .tbm at all -- but the guard was sprite.is_none(), a symptom test of the same shape as the one they say fixed their symptom not their cause. Now role == primitive. Six verdicts identical, 16 screens validate. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
39209eab05 |
port: the title's sweeps loop, the black hold is 9 units, and one claim refuted
THREE THINGS FROM THE DECODER, one of which I am not taking. REFUTED: "the developer splash is one composited quad, the bounding box of the three logos". The observed quad is 525x259 at (378,155). The three logos' bounding box is 500x421 at (390,164) -- a 259-tall quad CANNOT contain them, and palogo_anima alone starts at y=449, thirty-five pixels below that quad's bottom edge. The observed quad matches the union of gamearts_eff and seta_eff, 521x261 at (379,154), to about four pixels in every dimension -- and both of those are TRANSIENTS my own census flagged, dark by t=45, so a frame containing that quad is a build-in frame rather than the settled screen. I cannot see their draw stream, so I sent the arithmetic rather than a verdict, and the port keeps drawing three: I will not stop drawing an element on a claim whose stated identification excludes that element from its own bounding box. THE BLACK HOLD IS 9 UNITS, NOT 12. I authored 12 from Q7's luminance plateau of 0.17-0.23 s, supported by the menus' transition quad. The Decoder counted SUBMITTED QUADS instead -- luminance cannot separate the outgoing fade's tail from true black. Four frames with no sprite quad at all, at 2.284 units/frame derived from the disc as its own clock, gives 9.1 units = 0.152 s (6.9-11.4). That overlaps the luminance figure only at the top, and the true black is SHORTER still since both boundary frames carry picture. My 12 was supported by analogy -- a different screen's quad on a different path -- and a number that fits by analogy loses to one measured in place. verify-dwell's bound moved with it; both screens still agree. THE TITLE'S SWEEPS LOOP. The oracle shows the quad oscillating over its whole x range and resetting hard, one reset in the first title dwell and two in the second. The loop-length field could NOT have settled it, correcting a hope I had stated: both records declare exactly their last keyframe time, slack zero, and "loops at 600" and "runs once for 600 and stops" write the identical header. Verified on the two sweeps' LCM, since their periods differ: 600 and 720 realign at 3600 units, mean diff 0, against 0.438 at half that. Scoped to the title. The menus declare the same lengths but the oracle measurement is of the title, and my own weak evidence points the other way there -- best match with the sweeps off-screen, three times worse mid-screen, against a 73% on-screen duty cycle if they looped. Two weak signals in opposite directions is a reason to scope, not to pick. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
5dbc9aeac0 |
port: delete exit_ramp_units, invert the format's own rule, and guard a scale-0 leaf
FOUR THINGS, and the first is what MISSION section 3 calls the measure of
progress.
DELETED `exit_ramp_units` AND `exit_ramp_seconds`. They were authored because the
disc had no time slot on a group's final keyframe, so the ramp into it was the
one unknown duration per screen. Under the corrected record layout that keyframe
does not exist -- a group is an 8-byte header then frames x {u32 time; 36-byte
pose} and every pose is timed. VERIFIED DEAD BEFORE DELETING: setting it to 9999
(166 s) moved the boot's transitions by 0.04 s, which is wall-clock jitter, and
both uses in ScreenView are gated on a condition that no longer fires on any of
the export's 866 keyframes.
INVERTED THE FORMAT'S OWN RULE. `check.rs` enforced "the final keyframe has no
`t`; the disc has no time slot there" and FORMAT.md stated it. Both are now
backwards, and the validator fired 150 times on a re-export. I had not run
`check` between pinning the tag and measuring against the oracle -- the pixel
harness was green while the format validator was failing on every screen with a
multi-keyframe group. A correctness harness does not replace a format one; they
fail at different layers.
GUARDED A SCALE-0 LEAF, which the Decoder hit in its own renderer: its leaf
branch marked the element drawn unconditionally while the blit returned early on
zero scale, so a scale-0 leaf suppressed its parent and blanked the element --
live on all four loading screens. This port did not have the bug only because
authored/rendering.json happens not to list pgloading_loop5. That is an accident
of a gate written for another reason, not a defence, so `_draw_leaf` now reports
whether it drew and `_draw` falls back to the parent.
ISOLATED THE PACING QUESTION rather than leaving it as a suspected regression.
Legacy association: publisher 4.70 agrees, developer 3.92 DIFFERS. Corrected:
publisher 4.26 DIFFERS, developer 3.62 agrees. Both misses are ~0.03 s outside a
composite bound. The association traded which screen is marginally out; it did
not regress the pacing.
Bumped the pin c -> d for the parser and audio changes. Its headline renderer
change does not reach this port: sylpheed-cli builds from the workspace crate.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF
|
||
|
|
8994ca7c59 |
port: the 11.5 px was the fit's resolution -- and the lesson inverts
The Decoder closed it by ADDING OBSERVABLES, not by tuning. The vertex buffer carries positions and colours at the same instant, so all four quantities must agree on one t: quad A x solves to 357.88 and quad B x to 357.58, both +/-0.12 units, against 355.75 +/-1.54 and 354.09 +/-1.89 from the alphas. Alpha moves only 0.27-0.33 levels per unit, so one byte of quantisation is worth 1.5-1.9 units -- 6-8 px of sweep at 4 px/unit. That is the whole of the 11.5 px. At t=357.7 the centres land within 0.70 px and both alphas inside one level. THE LESSON IS THE EARLIER ONE INVERTED AND IT IS THE HALF WORTH KEEPING. Checking a wrong rule against alpha made it look confirmed; here the same insensitivity MANUFACTURED a residual that did not exist. An insensitive quantity does not merely fail to falsify -- it invents error. Solve on the fastest-moving field, check the slow one, never the reverse. I was already looking for a pivot rule to explain 11.5 px when they wrote; there was nothing to find. REFUTATION ATTEMPT, survived with a nuance: they state the leaf pivot is (200,90) on a 399x180 sprite, "the pivot is the centre, so rotation displaces it by nothing". Checked against my export -- pivot [200,90], sprite 399x180, true centre 199.5,90. It survives, but the sprite is ODD-WIDTH so the pivot is the centre to within half a pixel rather than exactly. No consequence against their 0.70 px agreement; worth stating because "displaces it by nothing" is the kind of sentence that later gets leaned on for a sub-pixel claim. verify-capture now poses the title at t=357.7 rather than 355: RMSE 21.07 -> 20.92, differing 1.82% -> 1.81%. Marginal, and it is the right pose for a stated reason rather than a better number. AND ptlogo_eff2 IS WITHHELD FOR A BETTER REASON THAN MINE. I had it on caution about untested generalisation; the Decoder points out it is on title_jp and MISSION section 7 scopes out "localisation beyond English", so it is not a question this port has to answer and the parked Japanese capture does not need reviving for it. authored/rendering.json now gives scope first and undecidability second. Widening scope to close a residual would have been the wrong trade. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF |
||
|
|
42474642e0 |
port: pin formats-pin-2026-08-29c -- the knob I tested last iteration was retired
I tested the wrong switch. SYLPHEED_KF_TIME_SHIFT is a superseded partial fix: it
got the association right but LEFT POSE 0 UNTIMED, which is exactly why the
untimed keyframe appeared to move from last to first. The real correction is the
DEFAULT in the tagged crate, with the old reading behind SYLPHEED_KF_TIME_LEGACY.
So last iteration's five rows measured a mismatch against a knob nobody should
use -- I suspected they were not decisive, I did not suspect the knob was retired.
THE CONSEQUENCE IS MUCH SMALLER THAN I BUDGETED. A placement group is an 8-byte
header then frames x {u32 time; 36-byte pose}, so pose 0's time is the group's
lead-in word and every pose is timed. Measured on the re-export: 866 keyframes,
0 untimed. `pose_at`'s "the final keyframe carries no t, so give it a synthetic
time" premise does not invert, it DISAPPEARS -- dead code rather than wrong code,
which is why nothing needed re-deriving. And the leaf now reads t=0 x=-639,
t=150 x=-39, t=540 x=1521, giving x=781 at t=355: the Decoder's predicted
top-left, and the 1300 px discrepancy is gone.
Pinned by tag, which is what MISSION section 2's tagging rule is for. BLOCKED was
wrong in both directions -- "cannot be taken yet" AND "only when that branch lands
on main". It arrives when the tag is pinned.
COST STATED: sylpheed-cli builds from the workspace crate, so until this reaches
main the exporter and the reference renderer read different decoders and
verify-screen compares two eras. verify-capture is unaffected -- it compares
against oracle captures and never touches the CLI. Revert to the path dependency
when the tag is an ancestor of main.
Oracle: publisher_logo 1.00% -> 0.75%, developer_logos 0.39% -> 0.33%, and
extras' differing region COLLAPSING from 736x525 to 398x295 at the sweep position
-- the residual localised onto the one element still in question. title unchanged
at 1.82%, now posed at t=355, the Decoder's FITTED sweep time. t=390 measures
1.65% and picking it would be fitting the pose to the score.
REFUTED, MINE: "ptlogo_eff2 is the single drawn element at a scale that is not a
whole multiple of 100%". That census was parents-only; the 45 leaves hold
thirteen distinct non-whole-multiple scales and 125% is among the rarest at two.
The claim's real content was "the only one the port draws" -- about my element
set, not the disc.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF
|