# Method traps already paid for Each line cost an iteration at least once. They are general β€” they are not about Sylpheed, they are about how this kind of measurement goes wrong. Like [`REFUTED.md`](REFUTED.md), this list had been living in the autonomous agent's loop prompt, i.e. nowhere durable. See [`README.md`](README.md) for the βœ…/🟑/❔ confidence convention itself. ## Controls * **Every result needs a control. A control that fails kills the instrument.** * **Run the known-positive through a new filter FIRST.** Three filters have been killed by their own control. When one fails, **read the known-good's disassembly** before assuming a shape. * **A measured negative is a result** β€” but a negative is only as strong as the route you ran, so **state its reach**. * **A null result needs its cause shown to have happened.** * **A result with NO unknowns is suspicious.** * **Census the whole set; always run the other population as the control.** **Zero partials is stronger than a majority.** * **A 2Γ—2 partition is the sharpest general tool** β€” both off-diagonals empty is a law. * **Re-derive a doc's own numbers as the control.** ## Inference * **Never conclude from ONE sample.** * **A law proved on one population is a hypothesis on the next.** * **Finding one exception does not imply a family.** * **Consistency is not proof. A suggestive coincidence is a coincidence until measured.** **An analogy is not a measurement.** * **Same layout β‰  same instance.** **Same record-name set β‰  same object.** * **A marker is only proven by what it leaves out.** * **A high-confidence SCORE is not a high-confidence MECHANISM.** * **Knowing HOW MANY is not knowing WHICH.** * **Round numbers matching is weak evidence β€” unless you read the constant.** * **My own last-turn result is a hypothesis too.** * **A global partition can understate a per-owner one.** * **A residual is measured against a population β€” name it.** ## Searching and tooling * **A search that returns thousands has no power; state the reach.** * **A substring match is not a hit.** **A regex miss looks like a null result β€” print one raw sample before believing a zero.** * **A derived table can be a cross product β€” measure its shape first.** * **After refuting an instrument, sweep everything that depended on it.** * **Before measuring how wrong a tool is, read what the tool actually does.** * **The instrument must pass its own control.** * **Classify a bulk before mining it. The residual is the prize.** * **Rank by similarity β€” the cliff is the finding.** But **read the values before trusting the rank.** * **Grep the nouns before designing the experiment β€” and believe it.** * **Grep gives you a file list β€” read *every* file on it.** * **The answer is often already in the doc that owns the subject β€” read it end to end.** A 🟑 often names its own route. * **Before re-trying a blocked idea, check whether the blocker's own doc already tried it.** * **Ship a regenerator with every artefact.** An artefact that moves by a pure reorder is a tool bug. * **Never print per-entry lines from a disc-wide sweep β€” aggregate.** ## Reading the data * **Read what a loader NAMES, not where it stores.** * **A field the disc never values still gets named by the loader.** * **An indexed read beats a deduped-pool adjacency read.** * **Check the whole string set, not the one matching word.** * **A dict keyed by record name across a multi-entry pak is a lie.** * **A set-difference over names hides reuse β€” join per USER.** * **A self-index names records, not files.** * **Case-insensitive hashing means two spellings can be one entry.** * **An "unresolved" name may be the wrong kind, namespace or prefix β€” or part of a cut asset.** * **A garbled value may be a real string in another encoding.** * **Two of my own counts disagreeing is a grammar clue.** * **A game's own typo is a join key.** * **A bias constant in the code is a join key.** * **A prefix trap: enumerate maximal `[A-Za-z0-9_]` runs, not `startswith`.** * **Re-deriving a format is not a finding β€” asking whether its values *resolve* is.** ## Mechanics that have bitten * **Never hand-convert a decimal VA β€” print `hex()`.** * **`grep -c` counts LINES** β€” use `grep -o | wc -l`. * **`Counter.most_common()` tie-breaks by insertion order β€” use `sorted()`.** * **Raw grep cannot see inside compressed pak entries.** * **Commit messages go in a file** (`git commit -F`); a literal `|` in a table cell needs escaping; `git log --all -- ` can hang. ## Runtime / emulator * **Look at the PNG** β€” and check its dimensions. * **"Animating" is not "still in a mission".** * **Dedup entity enumerations by position value.** * **Do not diagnose timing or liveness under gdb.** `ps %cpu` is cumulative. * **Classify screens by whole-image statistics, not named pixels** β€” a named pixel is only valid while the image sits at a known place, and nothing errors when it moves. * **Do not poll faster than the guest updates** β€” it manufactures a clean curve out of noise. * **A probe that never performs the action will "prove" the action does not exist.** * **The container's Canary binary can be older than the Canary source tree, and the failure mode is a hang, not an error.** After a merge into `sylpheed-re` the prebuilt `xenia_canary` had no `log_ui_draws`, no `mem_watch`, no `create_profile_if_none` β€” and an unknown cvar makes xenia open an SDL message box before logging is up, which headless is an unexplained freeze. Check before trusting a harness flag: `nm -C | grep cvars::`, and `build-canary Release` if it is missing. * **A capture armed *at* a screen only ever sees the steady state.** Anything about how a screen is built or animated has to be armed *before* it exists. Re-arming every few seconds and keeping every log tiles the approach: each F10 opens a new numbered file and closes the previous one complete. * **Measure animation in submitted frames, not in seconds.** `VdSwap` counts are the guest's own frames, so an emulator at 80 % of real time does not move them; a stopwatch reading does, silently and by an unknown factor. * **~~This game's menus drop d-pad presses shorter than ~0.3 s.~~ WITHDRAWN 2026-08-28 β€” the menu WRAPS, and I had not measured that.** The claim came from reading a cursor that ended up "one item short"; once wrap-around at both ends was measured ([`menu-navigation-semantics.md`](menu-navigation-semantics.md)), every one of those press counts is exactly right β€” four presses at 0.12 s moved four steps *through the bottom*, which lands one above where a non-wrapping menu would put it. **No press was ever dropped.** The real lesson is the general one: *a step count is only readable once you know the topology*, and I invented a hardware-flakiness story rather than testing the ends of the list. Still true and worth keeping: screenshot after every step and read the cursor, rather than trusting arithmetic over the press count. * **A screenshot taken right after a transition can catch a screen mid-fade.** A grab 2.5 s after β’· returned the game to the title showed the title art with no `PRESS β’Ά BUTTON` plate; one second later the plate was there. That very nearly went into the corpus as "the returned title has no plate". Sample a changing screen several times before writing down what it does *not* contain. * **Do not identify a menu cursor by label brightness.** The obvious oracle β€” "the focused label is the brightest row" β€” fails on this game's menus, because the background art is brighter behind some rows than the highlight is. It confidently named the wrong item on a frame whose ring was plainly elsewhere. Detect the **focus ring** in the gutter left of the labels instead (`tools/re-capture/menu_focus.py`, 254 vs <82 β€” no threshold tuning needed), and look at the PNG before believing either. * **`screenshot` samples at 0.5 Hz β€” it cannot time an animation.** Measured: ~2 s per grab (an `import` of the root plus an ImageMagick crop). A 0.4 s fade falls entirely between two samples, which is why a 40-frame burst across a screen change looked like an instant cut. For anything timed, record the display instead: `ffmpeg -f x11grab -framerate 30 -video_size x -i :98+, -t `, then read per-frame statistics off the file. Take the geometry from `xwininfo -root -tree`, the same way `bin/screenshot` does. * **A screen's brightness curve is not its fade quad.** The incoming screen's own elements animate in *after* the transition quad has cleared, so mean luminance keeps rising long after the fade is over β€” 1.47 s against a declared 0.97 s on one screen. Time the fade from where the frame is *pure black*, and take the ramp itself from the keyframes. * **A `sound.pak` entry is not a bank.** Entry windows straddle bank boundaries (the leading-region effect), so a naive "walk the RIFFs in this entry" gives the tail of the previous bank, then one of your own, then the *next* bank's header. It looks like a bank with the wrong number of waves. The tell is a `BANK id=` that is not the entry's own number, and the fix is the `seek` packet counts β€” they pin each wave's join exactly, so a straddle can be realigned rather than thrown away. 4 of 32 BGM banks look like counterexamples until this is applied. * **Do not time a video whose start you did not see.** An attract cycle sampled from partway through measured "β‰ˆ85 s", which was used to argue it was not the 137 s `ADV.wmv`. It was `ADV.wmv`; the sampling began 39 s in. Identify a video by **content** β€” match frames against the candidates and check the matched playhead advances monotonically β€” before reasoning about its duration. * **A near-black frame has no signature.** In frame-matching, low-contrast frames drive every correlation toward zero and the ranking becomes noise; one control frame scored 0.000 and another tied the wrong movie at 0.98. Read a dark frame's runner-up as *no evidence*, not as evidence for the runner-up. * **A screen that never changes is not necessarily hung β€” it may be a menu.** A 90-second screenshot loop that presses nothing will report a perfectly healthy `DIFFICULTY` menu as a standing hang, and that reading survived in this corpus for months. Before calling a static screen a hang, *look at it*, and press something. * **Never run `ps -ef` in this container.** Each of the three long-lived processes carries the whole loop prompt as its argv, so the listing is tens of kilobytes of the prompt you already have. Use `pgrep -x ` or `ps -o pid=,stat= -C `. * **A file extension is not a format.** `sound.pak` is full of `.slb` banks and names a `.xgs` project, and neither is an XACT container: the whole 1.08 GB archive contains zero `XGSF`, `SDBK` or `WBND` magic, and the executable has no XACT string at all. The names came from the authoring tool and survived into the shipped data. Check for the magic before committing to a parser for the format the extension advertises. * **A log-polling probe must compare a COUNT, not test for presence.** Timing an operation by waiting for its "done" line in a log gives an instant false positive if any earlier run left one there. This produced a confident "39.5 fps" that was really 0.1 s of matching a stale line β€” and it disagreed with the true 28.5 fps by enough to have flipped a conclusion. Snapshot `log.count(marker)` before starting and wait for it to *increase*. * **This container kills long emulator runs.** Twice in one session the `xenia_canary` process vanished with **no crash line in its own log** β€” once ~50 s into a boot, once ~145 s in at 1.15 GB RSS β€” while the box sat at ~1 GB free with swap exhausted. Budget dynamic experiments to fit in ~2 minutes of guest time, take the measurement early, and never read a disappearance as a guest crash without checking the log tail for one. * **When a frame-match will not resolve, sample denser before doubting the hypothesis.** Identifying a movie from 5-second screenshots gave contrast-23 frames, a playhead that would not join up, and a runner-up winning one frame β€” it looked like weak evidence for the wrong film. The same question sampled every 0.5 s (`ffmpeg x11grab` at 10 fps) answered at **0.999 with a strictly monotone playhead**. Sparse sampling does not weaken a signal gracefully; it turns it into noise that looks like a different answer. * **Check what instrumentation the local build already has before declaring a question blocked on tooling.** Two Q8 residuals were written off as needing "an emulator whose audio path can be observed". The emulator in this container already had `xma_param_probe` β€” visible in the startup CONFIG DUMP the whole time β€” added by this project for exactly that purpose. Read the cvar dump, or `nm -C | grep cvars::`, before concluding you cannot measure something. * **"Strings a function references" is not "arguments a function passes".** A loose filter over the strings in a caller produced a tidy-looking screen vocabulary; recovering the actual register argument at each call site removed half of it, including two names that appeared to be corroborated by independent measurement. If the claim is about an argument, simulate the registers β€” the cheap proxy manufactures agreement. * **A backward scan for `addi rN, …` misses arguments staged through another register.** PPC code often builds a pointer in a scratch register and moves it into place later (`addi r27, r11, X` … `mr r5, r27`). A backward search for the destination register alone found 7 of 48 sites and missed the two known-good ones. Simulate forward over a window instead, tracking `lis`/`addi`/`mr`. * **Check a measurement against the WHOLE reference set, not the part you were reading.** Two observed BGM stream sizes were recorded as matching "no bank's declared waves", which spawned a wrong theory that the decoder gets a window rather than a whole wave. They match `BGM_103` exactly β€” the check had covered only the `BGM_0xx` rows of the census, because that is the block that had been on screen. * **`instructions.function` is unpopulated for most rows in `sylpheed.db`.** A query keyed on it returns *nothing* rather than erroring β€” a `bctr` search scoped that way reported "no jump tables" for a function with two known ones. Scope by `address between and ` instead, and gate any such query on a function whose answer you already know. * **Before testing "state X behaves like this", check that your route actually reaches state X.** A test of "does β’Ά work in phase 4 state 0" was run by pressing β’·, on the strength of a decoded `4 β†’ state 0` edge β€” but the same basic block also writes the *phase* field, sending it to phase 2. The test probed the wrong state and produced a confident refutation of a live hypothesis. When a transition sets more than one field, read the whole block, not the one store you were looking for. * **A default filter can hide a whole screen, and the corpus will record it as absent.** `screen list`/`render` filter on `is_build`, which needs a `.rat` child; the developer splash has none, so it never appeared β€” and a page ended up stating it "cannot be rendered at all". It renders fine under `--all`. When two pages disagree about whether something exists, run the tool both ways before believing either. * **A logo that appears twice may be two different things.** The boot shows `SQUARE ENIX` at ~2 s and again at ~10 s; the first is the splash screen, the second is the intro movie's own opening frame. Reading the capture alone gives "three logo screens". Cross-check a repeat against what is *supposed* to be playing at that moment before counting it as its own screen. * **Do not hard-code an expected element count into a structure scan.** Looking for a `.rat` record's keyframes, I required 29 strictly-increasing times because a header word read `30`. The records hold **three** keyframes in exactly the layout I was testing for, and the over-constrained filter reported the layout as refuted. Scan for the *shape* (here: `0x??ffffff` fade words at a constant stride) and let the count fall out. * **Read the target text before writing a patch pattern for it.** Two iterations running, a scripted edit to `HANDOFF.md` failed its assertion because the paragraph had been rewrapped since the pattern was written β€” and both times the finding landed in `docs/re/` while the page the port agent reads kept the stale version. The commit's file count is the tell: if it is one short, an edit was silently dropped. * **Two renderers agreeing is not evidence the field is right.** The port's composite and `sylpheed-cli screen render` matched on `main_menu` to 3/255 β€” and *both* omitted two elements the game draws, because both read one field through one decoder. Agreement between consumers of the same decode measures the decode's determinism, not its truth. Only a framebuffer capture is an independent check. (Reported by the port agent, 2026-08-28, after a capture crop found the missing bracket.) * **`./target/debug/` is STALE in this container β€” `CARGO_TARGET_DIR` is redirected.** The Dockerfile sets `CARGO_TARGET_DIR=/sylph-home/re/target-container` so a fresh `cargo build` writes there, while `./target/debug/sylpheed-cli` remains whatever was built on the host, hours old. A decoder fix was verified against the stale binary and produced byte-identical "before" and "after" renders β€” which reads exactly like "the change had no effect" rather than "you ran the old code". Run the binary cargo actually wrote, or `cargo run`. * **A pivot-equality test conflates rounding with real breakage.** Checking `pivot == texture/2` over `GP_TITLE` flagged 74 % as mismatched, which reads as a systemic decode failure. Split by magnitude it is 17 exact, 54 off by ≀ 1 px (odd dimensions rounding), and only **24 off by > 8 px** β€” and those 24 are the real finding, concentrated on one element family. Bucket a mismatch by size before reporting a rate. * **I have now broken the "read the target text first" rule three times.** Each time a scripted `HANDOFF.md` edit asserted on a paragraph that had been rewrapped by an earlier edit in the same session, and each time the finding landed in `docs/re/` while the port's page kept the stale text. Writing the rule down was not enough; the working fix is mechanical β€” `grep`/`sed -n` the target paragraph in the same command that patches it, and check the commit's file count. * **"The nearest header after the name" is not a nameβ†’header mapping.** Looking for each sprite's `T8aD` header by searching for its name and taking the next `T8aD` returned *the same header for every sprite* β€” all the names live together in the declaration table near the top of the bundle, so "next" is always the first one in the file. It failed its own control instantly: different sprites reported identical dimensions. Match on something the header itself carries (here width Γ— height against the decoded PNGs) rather than on proximity. * **`cargo build` passing does not mean `cargo test` compiles.** Adding a field to `T8adImage` built the library fine in 1.48 s β€” and broke two *test-only* struct literals, so `cargo test` failed to compile with exit 101. A green build on the binary is not a gate for a struct change; run the tests before believing a data-structure edit is safe. * **Detach long background work with `setsid`, or a tool-call timeout kills it.** Three attempts to run the disc-gated test suite in the background died at exit 143 with an empty log β€” the timeout's SIGTERM reached the whole process group. `setsid nohup cargo test … &` survives, and the run then completes across later tool calls. An empty log from a "background" job usually means it was killed, not that it produced nothing. * **Validate a recommendation before leaving it as advice.** Having concluded a question was undecodable from the disc, I named a per-draw GPU capture as the next route "because it reads the blend state". It does not β€” checking the capture's own source showed it records textures and vertex attributes and no blend registers. A named next step is a claim like any other. * **Name the element from the draw's own coordinates before building on it.** A skewed draw in a title capture was called "the swoosh" because the swoosh is the screen's only diagonal element β€” an identification by elimination, never checked against the draw's NDC extent. It was wrong: the quads span the full screen and beyond, the swoosh is a 234-pixel band. Five iterations of analysis pointed at the wrong element on the strength of it. Converting the coordinates takes one line and should come first. * **A census over a bundle's top-level table is not a census over the bundle.** The keyframe rotation field read "always zero on this screen" for several iterations because every scan walked `GP_TITLE` build 4's **declaration table**, and the rotated elements are **nested leaf records** reached through an `opt ` link. The bytes were right; the *reach* was wrong, and a negative stated without its reach reads like a fact about the disc. Say which region a negative covers, and check whether the thing you are looking for lives outside it. * **Do not assume 4-byte alignment when scanning raw bundle bytes.** A nested `RATC` blob starts wherever the parent's chunk stream leaves it β€” `ptloop01.rat` sits at `0xbb5966` β€” so its 40-byte keyframe blocks are odd-aligned. A scanner that filtered candidates on `%4 == 0` found **0/3** of its own control blocks and under-counted the corpus by **16 341** blocks, all of them nested. It cost nothing to catch, because the control was in the script. * **`pkill -f ` matches the shell running it.** `pkill -f kfscan.py` from a `bash -c` whose command line contains `kfscan.py` kills its own shell: the tool call returns exit 143/144 and the edit that was queued behind it never runs. The same trap makes `pgrep -f` self-report β€” a "still running? yes" that was the shell seeing itself, on a job that had already finished. Use `pgrep -x`, or match on a path the current command line does not contain. (This is the second `ps`/`pgrep` entry in this file; the first is about `ps -ef` dumping the loop prompt.) * **Grep a legend and you count the legend.** A sweep for elements with a rotation reported "1 element" in every build on the disc β€” the header line `a=alpha r=rotΒ°` matched the ` r=` pattern. A uniform count across heterogeneous inputs is the tell. Make the pattern require the *value* (` r=-?[0-9]+`), and sanity-check that a known-negative build reports zero. * **A pattern over the archives you happened to check is not a negative.** I wrote "rotation appears to live only in nested leaf records", correctly scoped to the three paks I had swept β€” and the sweep refuted it on the fourth, within the hour, because the alphabet had not reached `GP_DIALOG` yet. Scoping the claim was what made it cheap to withdraw, but the better move is to let the sweep finish before writing the sentence at all: a negative that is still running is not a finding, it is a prediction. * **A field that is implemented is not a field that is measured.** The pivot-anchored scale term had been in `blit` for months and reasoned about in the handoff, and no capture had ever exercised it, because every element anyone had looked at was at 100 % scale where the term is identically zero. Look for the case that makes a term *large* β€” here 600 % and 800 % scale, worth 450 px β€” and check it there. A term you cannot distinguish from zero has not been verified by any amount of agreement. * **Printing a series' minimum and reading it as its range.** I summarised a captured alpha series as "constant Ξ± β‰ˆ 33" and built a contradiction on it β€” the summary printed `min_alpha` and no maximum, and the series actually ramps 34 β†’ 255 β†’ 33. The tell was there in the same table: the column beside it said *14 distinct colours*, which a constant series cannot have. When a summary statistic and a distinct-value count disagree, the summary is wrong. * **Calibrate on one element, test on another.** Fitting a declared ramp to a capture has two free parameters (rate and offset) and will "succeed" against almost anything β€” my first attempt scored RMS 128/255 and I nearly read the numbers rather than noticing the search could not reach the ramp at all. The version that means something: fix the clock from element A's ramp, check that fix against A's own next landmark, then apply it to element B in the same frames with **nothing left to tune**. That is what turned "the shapes look different" into "still at 255 nine frames after it should read 32". * **Two measurements can both be right and still disagree, when a heuristic sits between them.** A change to the keyframe time association is favoured 26Γ— by an animation measurement and rejected by a static-render comparison. Both are sound: `rest()` picks a pose by *longest dwell*, which is a heuristic layered on top of the times, so moving the times moves its tie-breaks. The mistake would be to treat one as refuting the other and drop the loser. Name what each measurement actually constrains β€” here *timing* versus *pose selection* β€” before deciding they conflict. * **Before believing a render diff refutes a decode, find out which element moved.** I reported a 13 % render change as evidence against a keyframe-time hypothesis, on a brightness argument about language twins. One element accounted for all of it, and that element is a transient flash with **no resting pose** β€” so the difference was our own `rest()` heuristic guessing, and it would have guessed under either reading. A whole-image statistic (pixels changed, mean luminance, correlation) tells you *that* something moved, never *what*, and an argument built on it can be entirely about your own code. Localise the diff to an element and read its keyframes first. * **A fallback that only runs when its precondition fails is not a fallback.** `rest()` picks the longest-dwell keyframe when no plateau exists β€” but a dwell gap is time spent *moving between* two poses, and the only case where an endpoint is genuinely held is when the two poses are equal, which is exactly the plateau the first path already caught. So the fallback is guaranteed to be guessing every time it is reached. When a heuristic's justification is the same condition that routes around it, it has no justification. * **Siblings adjudicate a rule that a single element cannot.** Deciding what `rest()` should return for an element with no held pose looked like it needed a capture we do not have. It did not: the developer splash carries three glows with identical structure and identical times, differing in one alpha byte, and any rule that makes one of them behave completely differently from the other two is producing an artefact. Parallel constructions in the same bundle are a free control β€” look for them before concluding a question needs new measurement. * **Backgrounding with `&` in a compound command loses the working directory.** `cd X && cmd &` followed by more work in the same call left the shell reset, and a heredoc two lines later wrote nowhere β€” while the `echo` after it still printed "written". A success message after a failed redirect is not a success. Use `(setsid nohup … &)` and keep the file-writing in its own call. * **Read the whole function before describing what it does.** I wrote "a zero-scale pose paints a 1-pixel sliver" into a tool's docstring, from the `.max(1)` in the size arithmetic. Two lines above it, a guard coerced `scale == 0` to **100 %**, so the `.max(1)` never saw a zero and the element rendered at *full size* β€” a much larger bug than the one I described. I had read the lines I went looking for and stopped. * **Never diff a file a background job might still be writing.** A comparison loop over an output directory reported `GP_TITLE` build 6 as changed; the pixel diff was then all zeros, because `cmp` had read a half-written PNG. Compare only after the job signals done, or skip the newest file. A "difference" that vanishes on re-examination is usually a race, not a flake. * **A `screen list` BUILD index is not a pak ENTRY index.** `GP_TITLE` happens to map 1:1 (16 builds, entries 0–15), which is how the assumption survived; `GP_DIALOG` does not β€” its build 0 is **entry 2**. Indexing a pak directly with a build number silently reads a different bundle and reports a confident wrong answer: a census control asked for the two rotated elements of `GP_DIALOG` build 0 and got zero, from a bundle that genuinely has none. `screen list` prints the mapping; use it rather than assuming identity. * **"There is no flag" is not "there is no way".** I declared a Japanese-locale capture impossible in this container because `user_language` is `DECLARE`d and never `DEFINE`d, so no cvar exists to pass. That was true and not the question. The language is *persisted*: canary builds `XConfig` over `/xconfig.settings`, `SetDefaults()` only supplies the value when the file has none, and the file is writable. One grep further β€” for where the setting is **stored** rather than where it is **configured** β€” turned a "blocked, needs a human decision" into a two-line edit. When you conclude a capability is absent, check the persistence layer before writing it up. * **Locate a binary field from struct landmarks, not from a hard-coded offset.** `user.language` was found by scanning for `music_volume`'s `0.7f` (a known default at a known relative offset), then cross-checking `language == 1` and `country == 103` against the emulator's own `SetDefaults()`. Three agreeing landmarks, no offset assumed, and the check re-runs every invocation β€” so the tool fails loudly if the layout moves instead of silently patching a neighbouring field. * **A stale oracle reports the world, not itself.** A 787 s run concluded "the game never reached the title". It had, repeatedly: `wait_title.sh` was still sampling the single pixel (625,618) that `is_title.py` had already been written to replace β€” a 1280Γ—720 coordinate against a 1279Γ—675 game surface, always reading the copyright line. The replacement existed *in the same directory*, with a docstring naming the bug. Before believing a negative from a probe, check whether something in the toolkit already supersedes it, and run the probe against a known-positive capture first β€” `is_title.py` scores 753 on a real title frame and 327 on the main menu, which takes one command to confirm. * **Run the control before theorising about the difference.** A Japanese-locale run failed to reach the interactive title, and I wrote up the open question as "if English reaches it, the difference is the locale". English does not reach it either β€” 75 samples over 734 s, every one a miss. The locale had nothing to do with it, and two iterations framed the problem around a variable that was never implicated. The control cost one run and would have cost the same one iteration earlier. * **Log every sample, not just the hit.** The first probe printed nothing until it succeeded, so a 787 s failure was indistinguishable from a probe that never ran. Printing the oracle's value each tick turns "no title" into "75 samples, max glyph 0", which is a measurement β€” and it is what showed the black-screen and attract-movie phases were being sampled correctly all along. * **Know when to stop paying for a non-blocking answer.** A capture that would settle two 🟑 residuals β€” neither of which blocks the deliverable β€” absorbed five iterations. Each one produced a real finding (a broken oracle, a wrong "impossible", a locale red herring), so none was wasted, but the *question* never moved. The signal to stop is not "this is hard", it is "the thing this unblocks is not blocking anything". Write down the recipe you built and leave it for a session with a reason to spend the time. * **Audit what a renderer silently SKIPS, not only what it draws.** `compose` has half a dozen `continue` paths, and a sprite-name miss is one of them β€” an element vanishes with no error, which is exactly how the port agent's missing bracket happened. `screen render` already prints its omissions; reading that list for the five screens that matter took one command and turned up a full-screen opaque element dropped on three of them. Any composer with skip conditions should be asked to name what it skipped before it is trusted. * **An element's declared name is not necessarily its texture's name.** The `opt ` link on `pteff05.t32` leads to a `T8aD` registered as `8AX`, so a lookup by declared name misses. Before concluding a texture is absent, look at the bytes the link actually points at β€” the name in the declaration table and the name the sprite table keys on can differ. * **A fit whose residual is large for every model is a broken method, not a close call.** Comparing our composite to a capture pixel-wise gave mean abs errors of 10–14 for gamma, for a linear scale, for everything β€” and a non-monotonic transfer curve (render 96–127 mapping *brighter* than render 128–159). The cause was edge misalignment: at correlation 0.947 a bright pixel in one image routinely lands on a dark one in the other. Restricting to patches that are flat in **both** images dropped the residual to 0.2–1.1. When every candidate model fits badly, stop choosing between them and look at what the comparison is actually measuring. * **Say when a control failed to discriminate, rather than reporting it as a pass.** The held-out screen for a tone-curve fit was the developer splash, whose flat regions are pure black β€” every model scored β‰ˆ 0.00 error there. That is not corroboration; it is a test with no power, and reporting the 0.00 as agreement would have dressed an untested claim as a verified one. * **Time your probe against the thing you are probing for.** Four runs concluded "the game never reaches the title". `screenshot` costs **10.8 s while xenia is running** and **0.117 s once it is killed** β€” 92Γ— β€” so a two-grab polling loop samples every ~41 s, against a title screen this corpus documents as lasting a few seconds. The harness was blinking slower than the event. Before believing a negative from a polling loop, measure its interval and compare it to the duration of what you are waiting for; and measure the probe's cost *under the same load as the run*, because idle timing here was off by two orders of magnitude. * **Do not change a display setting and a capture behaviour in the same run.** `kernel_display_gamma_type = 0` brightens the frame, and `skip_intro.sh` classifies movie-vs-static on an *absolute* rmse threshold β€” so the gamma change biased the very classifier the run depended on. Harness thresholds tuned on one output configuration are not portable to another. * **Fixing the instrument is how you test the explanation that blamed it.** I diagnosed four failed runs as "the probe samples slower than the event", wrote it up, and withdrew three earlier conclusions on that basis. Building the faster probe refuted it: at 100Γ— the sampling density, over 420 unbroken seconds, the event still never occurred. The diagnosis was a real defect that happened not to be the cause β€” the most seductive kind, because fixing it feels like confirmation. A plausible mechanism that explains a failure is a hypothesis, and the fix is its experiment, not its proof. * **Ask whether the screen is drawing anything before explaining why a feature of it is missing.** Four iterations produced four explanations for "the title is not detected" β€” a stale oracle, the locale, the sampling rate, a one-shot boot window β€” and three of them were refuted in turn. Splitting a raw root grab into bands then showed the game surface at **0.08 % non-black, mean 0.07**: it was rendering black. The check costs one command and belongs at the *front* of that sequence, not after the fourth hypothesis. A detector reporting absence is ambiguous between "the feature is not there" and "nothing is there". * **A newly found fault does not retroactively explain older failures.** The black surface is real and reproducible *in the run where it was measured*. Earlier runs in the same session had measurable content (mean 33.1, mean 122.8, non-zero frame-to-frame rmse), so the failure mode changed over the session. Reaching back to re-explain earlier results with the newest cause would have been the fourth over-reach in a row. * **Validating a probe on static test images tests its ANALYSIS, not its ACQUISITION.** The fast probe's counter was controlled against committed captures and matched exactly (753 / 327), which proved the numpy expression right and the frame source untested. The source was the broken half: a long-lived x11grab stream degrades from 3.98 to 1.60 fps and then freezes, repeating one stale frame. The check that catches it is to read the same screen with an **independent grabber at the same moment** β€” `import` said 125.65 where the stream said 5.21. Do that during a run, not only at design time. * **A dense negative from one instrument is worth less than one cross-check.** "2 391 frames, zero hits" reads as overwhelming and collapses to nothing once the frames may all be the same frame. Sample count is not evidence of coverage unless the samples are known to be independent. * **A probe that cross-checks itself turns "no result" into a result.** Three successive "the title never appears" claims were withdrawn because the instrument was broken each time. The version that finally held prints its own reading beside an independent grabber's every 60 s β€” 9 checks, 8 agreeing to within 2 % β€” so the negative arrives with its own evidence that the frames were live. Building the cross-check into the tool costs a few lines and is what separates a measurement from a silence. * **A hedge in a code comment is an unmeasured claim.** `compose` said the derived paint order "reproduces both measured orders up to ties" β€” a sentence that sounds like a result and was neither measured nor kept current: there were three measured orders by then, not two. Measuring it took one example program and turned a hedge into a bounded number (exact on 4 of 5 bundles; the 5th off by 8 pairs, all ties). Grep your own comments for "up to", "roughly", "essentially" β€” each one is a claim nobody has checked. * **Count the cases that can actually bite, not the cases that match the pattern.** `EXTRAS` has 15 tied pairs in its derived paint order, which reads alarming. Only **2** of them overlap, and a tie between non-overlapping elements paints identically either way. Reporting 15 would have overstated the risk by 7Γ—; the useful number is the one filtered by whether the difference can reach a pixel. * **An exhaustive field search needs a positive control, or "found nothing" is worthless.** Scanning a header for a field that reproduces a measured ordering returned zero hits β€” which could equally mean the field is absent or the scan is broken. Running the same scan against a *known* ordering (declaration order) returned **64** hits, proving the scan finds ordering fields when they exist. Only then is the zero a finding. The control costs four extra lines and turns a silence into a negative with reach. * **A stale row on the handoff page is worse than a missing one.** The port was told SE audio was "undecodable from the disc" while the linked doc had already retracted that and located three cues decoding to PCM. The finding landed in `docs/re/` and the page the port reads kept the superseded text β€” the fourth time in this corpus. Writing the rule down has not worked, so there is now a tool: `tools/re-capture/handoff_lint.py` flags every HANDOFF line that makes a strong negative claim and links a doc containing retraction language. It found this row, and one benign false positive, on its first run. * **Narrow a risk before reporting it, then narrow it again.** "15 tied pairs on `EXTRAS`" became 2 (only overlapping ties can paint differently), then 1 (one of those two is a `loop*` animation the compositor skips), then "consistent with the capture". Each step is a filter that costs one query and changes the number by more than an order of magnitude in total. A raw count is almost never the number a consumer needs. * **Nothing was checking that the docs' cited evidence exists.** A sweep of every relative link under `docs/` found **16 broken**, and two of them were the figures backing the UI layout decode's headline claim β€” the port's foundation, unreachable from its own page, because a path in `structures/` was written one directory too shallow. Eleven were wrong relative depth with the target present; five name files that do not exist. Evidence that cannot be opened is not evidence, and a link is exactly the kind of thing no one re-reads. `tools/re-capture/doc_link_check.py` now does it, and also flags targets that resolve to a **zero-byte** file β€” which looks correct in every listing. * **Repair in bulk only when the counts pair.** The fixer rewrote 11 links; the checker went from 1 038 resolving / 16 missing to 1 049 / 5. +11 and βˆ’11 against 11 edits is the confirmation that the pass did what it said and touched nothing else. A bulk edit without that arithmetic is a hope. * **When two candidates carry the same content, compare their DIFFERENCE against the oracle, not themselves.** `8AX` (1280Γ—720) and `ptbase` (640Γ—360 at 200 %) are the same artwork, so no comparison of either against a capture can separate them β€” every such test had been read as "inconclusive, needs a per-draw capture". What separates them is `8AX βˆ’ upscale(ptbase)`: the detail only one of them has. Correlating the capture's residual against *that* answered it statically, on two screens, with matched controls. * **A ratio that saturates at 1 under noise is not a discriminator.** A pixel-pair test cleanly separated upscales (0.00–0.72) from native (0.98) and put the capture at 1.01 β€” apparently decisive. Additive noise raises both terms of the ratio equally and drives *any* value toward 1, and fitting a noise term showed both hypotheses reproduce the observed numbers. Before believing a ratio, ask what it does as noise grows. * **Read what a cvar DOES before building an experiment around it.** I planned a run to decide whether a measured gamma was "canary's BT.709 output stage", by setting `kernel_display_gamma_type = 0` and re-fitting. The cvar is not an output stage: it is the value a `kStub` **getter returns to the guest**, which the game uses to build its own ramp, and canary applies *that* ramp in the swap path. The experiment could never have isolated a stage that does not exist β€” and it would have perturbed the capture harness as a side effect. One grep for the cvar's definition replaced a planned emulator run with a better one. * **An absence in a log is only evidence if the log would have shown it.** No gamma lines appear in any run log here β€” but kernel exports log at Debug and this harness masks Kernel logging entirely (`log_mask = 13`), so their absence is guaranteed regardless of what the game did. Check the logging configuration before reading silence as a result. * **Check what an experiment actually needs before filing it behind a blocker.** The gamma question sat parked behind "needs the emulator to reach a menu" for several iterations. It needed the emulator only to **boot** β€” video init, and the `VdGetCurrentDisplayGamma` call, happen in the first seconds, long before any title screen. A blocker that stops one experiment does not stop every experiment in the same area, and it is worth re-reading the parked list against what each item really requires rather than against the area it belongs to. * **A default value is evidence.** Whether the game writes a gamma ramp looked like it needed a GPU trace. It mostly did not: canary initialises the ramp table to **identity** and applies it through a shader that is a pure LUT lookup with no other transfer. An unwritten ramp is therefore a no-op, and any non-identity transfer in the output implies a write. Reading what a field holds when nobody has touched it turns "I cannot observe the write" into "the write must have happened" β€” cheaper than instrumenting, though it stays an inference and should be labelled one. * **Name the weak joint of an inference in the same breath as the conclusion.** The chain above assumes our composite reproduces the pre-ramp framebuffer, which is the one step that could be wrong. Writing that down beside the conclusion β€” rather than only the supporting facts β€” is what lets a later reader attack it instead of inheriting it. * **The startup config dump is the config FILE, not the effective flags.** A run passed `--log_mask=12 --log_level=3` and its dump printed `log_mask = 0, log_level = 2` β€” while Kernel Debug logging was demonstrably on, which is how that run's finding was obtained. So a dump can neither confirm nor refute a command-line override, and reading one as "my flag was ignored" is a mistake I nearly made with `trace_gpu_stream`. Verify a flag by its *effect*, not by the dump. (This does not undo the earlier `user_language` conclusion: absence of a *name* from the dump still shows the cvar is unregistered.) * **`kill -9` destroys anything that finalises on shutdown.** Canary starts a GPU trace at init but only closes it in `GraphicsSystem::Shutdown()`, so the hard kills this session used routinely could never have produced a trace. Before concluding a feature is broken, check whether the way you stop the program is what discards its output. * **Bound a risky experiment from inside, not by watching it.** A boot-time GPU trace on a disk at 95 % could have filled it between two tool calls. The runner carried its own watchdog that killed the emulator the moment the output passed a 2 GiB cap, so the experiment was safe regardless of how coarsely I polled. The watchdog never fired, which is the point β€” it cost nothing and removed the need to gamble on timing. * **A cvar existing does not mean the feature is compiled in.** `trace_gpu_stream` parses, is documented, sets state, and does nothing: the code that opens the trace sits behind `#if XE_ENABLE_TRACE_WRITER_INSTRUMENTATION == 1`, which `trace_writer.h` ties to `#ifdef NDEBUG` β€” off in release. Two runs and two wrong hypotheses before reading the `#if`. When a switch produces no effect, follow the code from the switch to the output and look for a compile-time gate before theorising about runtime causes. * **Test a compile-time gate against the binary, with a control.** A string that exists only inside the guarded block (`_stream.xtr`) settles it in one command: 0 occurrences in the release binary, **1** in the debug binary. The debug build is the control that proves the test can find the string when it is there β€” without it, "0 occurrences" is just as consistent with a bad grep. * **`T8aD` headers sit in RATC child order β€” use that, not the size, to name one.** Attributing a sprite header to a sprite name by matching decoded dimensions fails whenever two sprites share a size, and `GP_TITLE` build 4 has such a pair (`ptlogo_back2eff` / `ptlogo_back2eff5`, both 1133Γ—280) that differ in exactly the flag being studied. Header order matches child order β€” verified 18/18 on that build against the dimensions as an independent check β€” so ordering resolves what size cannot. * **Identical summary statistics are not identical data.** Two same-sized sprites reported the same %opaque, %clear and mean alpha to one decimal, which reads as "the bundle stores this texture twice". Pixel-comparing them gives max abs diff **21**: two different renditions of one image. Three matching summaries over 300 000 pixels is weak evidence of identity; the comparison costs one line. * **A pattern that is perfect on one screen can be near-chance on the disc.** The `0x02` bit matched `eff` names 10/10 on `GP_TITLE` build 4, which is exactly the sample size that makes a rule feel found. Disc-wide it holds 46.8 % of the time β€” the build's artists simply named their effect sprites consistently. An 18-element bundle cannot distinguish a format rule from a local naming habit; before reporting an implication, run it over the corpus, not the example that suggested it. * **When an association survives a refuted implication, the counterexamples are the finding.** The bit is 3.3Γ— enriched for `eff` names and the sprites that break the rule are `pv_loading_ring0`, `pv_loading_light0`–`3`, `px_bunk_line` β€” rings, glows, lights, thin lines. That the exceptions are all effect-like artwork *without* the naming convention says more about the field than the rule it broke did. * **Report a classifier's lift over its base rate, not its accuracy.** The `0x02` bit's association with bright-RGB/low-alpha art gives a best single threshold of 76.5 % accuracy β€” which sounds like a finding until you notice that always guessing the majority class scores 64.1 %. The lift is 12 points, the distributions overlap, and it is a tendency rather than a rule. An accuracy quoted without its base rate is not interpretable. * **Park a field after N failed hypotheses, and say what was eliminated.** Four candidate meanings for one flag bit were each refuted at a cost of roughly an iteration apiece, and the bit blocks nothing measurable. Stopping is the right call, but only if the negative space is written down β€” additive blend, name, lifetime, premultiplied alpha β€” so the next attempt starts where this one ended rather than at the beginning. * **Nothing was checking that a doc's figures match its committed data.** A number is written once from a run and then lives in prose that gets edited around it; the data file beside it is regenerated independently. All 19 headline figures across four censuses do currently agree (`tools/re-capture/doc_figure_check.py`), which is worth knowing rather than assuming β€” but the checker had to be written **numerically**, because the first version grepped for the doc's formatting (`14 709` with a thin space, `33.7` rounded from `33.66`) and reported almost every figure as a mismatch. A consistency check that fails on formatting will train you to ignore it. * **A detached job you never check can outlive many iterations.** Two `setsid nohup cargo test … &` runs from earlier iterations were still alive **four hours** later, one child at 89 % CPU for 3 h 26 m, holding the load average at 14 on 12 cores. `setsid` was added precisely so a tool-call timeout could not kill them β€” which also means nothing kills them. Check that a backgrounded run actually exited before launching the next one, and prefer `ps -o etime=` over assuming. * **Know whether your verification gate can terminate.** `build-reborn test` includes a disc test that decodes 166 `.xpr` files (1.4 GB) and had not finished after 3 h 26 m of CPU. Every "green run" from it is partial unless the suite is stated to have terminated β€” so report the suite count and elapsed state rather than the word "green", which is what I should have been doing all session. * **"Cannot finish" and "takes an hour" are different claims β€” measure before choosing.** I wrote that the verification gate *cannot terminate*, from having watched it run 3 h 26 m. Timing the work directly: 19 of 166 containers exceed 25 s, `Stage_S02` completes in **144 s**, and one full pass is ~45–60 minutes. Nothing hangs. The 3 h 26 m was that hour of work at a load average of 9–14, inflated by my own duplicate runs. A slow thing observed under contention looks like a stuck thing, and the correction matters: an hour-scale gate can be run deliberately, a hung one cannot be run at all. * **Separate the question with evidence from the question tangled with it.** The keyframe-time reading was held back for several iterations by one objection: adopting it changes `rest()` on a single element. But `rest()` reaches the times only through a fallback that is unsound whenever it runs, and the shift does not fix that fallback either β€” so the objection was never about the times. Timing had three discriminating measurements; pose-selection had a heuristic guessing. When a conclusion is blocked, check whether the blocker is actually evidence about the same thing. * **A predicate over adjacent PAIRS silently misclassifies a one-element list.** "Has a plateau" was implemented as *any two adjacent keyframes share a pose* β€” which is false for a single-keyframe element, so 1 502 static elements were counted as having a *guessed* rest pose and the published defect rate was 65 % too high. The error only surfaced when acting on it: suppressing those elements dropped the title's correlation by 0.27, because they include the backgrounds. Whenever a rule quantifies over pairs, ask what it says about a list of one. * **Acting on a claim is a better test of it than re-reading it.** The single-keyframe flaw survived a disc-wide census, a write-up and a handoff row. It died the moment the rule was used to change a rendering, because the result was visibly worse. If a measurement implies an action, take the action on something you can score. * **Score a rule where it can actually differ, or you will measure nothing.** Three rest-pose rules rendered builds 4, 5 and 6 to *identical* correlations β€” not because they agree, but because the code they change is unreachable on those screens. The signal was on the two splashes, the only builds whose elements reach the fallback at all. Identical results across variants is a clue that the variant is not being exercised, not evidence that the choice does not matter. * **An argument from symmetry is a prediction, not a refutation.** I killed the last-keyframe rule because it treats one of three sibling glows differently, which felt like an artefact. Measured, it is the better rule on both screens where it applies. Aesthetic expectations about how authored data "should" look are worth stating as hypotheses and worth nothing as verdicts. * **A blank render is a correlation of NaN, not a low score β€” notice which.** Applying "rest = last keyframe" to every element made both splashes render with zero variance, so the correlation was undefined rather than poor. Reading the NaN as "the tool failed" would have hidden the result; it *was* the result, and the strongest form of it. Check for degenerate output before treating a missing number as a broken measurement. * **When a model predicts something and the measurement refuses, the model is usually incomplete rather than wrong.** The shifted time reading says when each pose is reached, which looked like it implied the final pose is the resting one. It does not: it says nothing about whether the group is *played to completion* while the screen is still up. The exit is the dismissal, so a displayed screen never reaches its last keyframe. Both claims survive; what was wrong was the step between them that I supplied. * **The control is what turns a helpful edit into a finding.** Suppressing the `_eff` glows raised both splashes from β‰ˆ0.96 to β‰ˆ0.998 β€” on its own that reads as "we over-draw glows". Running the same edit on the three screens that settle made them **worse** (βˆ’0.002, βˆ’0.092, βˆ’0.107), which is what makes the result mean something specific: the over-draw is confined to transient screens, where a plateau mid-animation is not evidence the element is on screen at rest. An edit that improves one set of cases is only interesting once you have shown it damages the cases where it should. * **When a model reproduces durations but not positions, the missing piece is an origin, not a rate.** Playing the splash timeline matched every element's on-screen *length* to within 2 % while placing the elements in the wrong part of the run. That pattern names the gap precisely β€” the rate and the shape are right, so what is absent is a per-group start β€” and it pointed straight at the one undecoded word in the group header. (Which was zero, but the diagnosis was still what made the next step obvious rather than a search.)