# Method traps already paid for Each line cost an iteration at least once. They are general — they are not about Sylpheed, they are about how this kind of measurement goes wrong. Like [`REFUTED.md`](REFUTED.md), this list had been living in the autonomous agent's loop prompt, i.e. nowhere durable. See [`README.md`](README.md) for the ✅/🟡/❔ confidence convention itself. ## Controls * **Every result needs a control. A control that fails kills the instrument.** * **Run the known-positive through a new filter FIRST.** Three filters have been killed by their own control. When one fails, **read the known-good's disassembly** before assuming a shape. * **A measured negative is a result** — but a negative is only as strong as the route you ran, so **state its reach**. * **A null result needs its cause shown to have happened.** * **A result with NO unknowns is suspicious.** * **Census the whole set; always run the other population as the control.** **Zero partials is stronger than a majority.** * **A 2×2 partition is the sharpest general tool** — both off-diagonals empty is a law. * **Re-derive a doc's own numbers as the control.** ## Inference * **Never conclude from ONE sample.** * **A law proved on one population is a hypothesis on the next.** * **Finding one exception does not imply a family.** * **Consistency is not proof. A suggestive coincidence is a coincidence until measured.** **An analogy is not a measurement.** * **Same layout ≠ same instance.** **Same record-name set ≠ same object.** * **A marker is only proven by what it leaves out.** * **A high-confidence SCORE is not a high-confidence MECHANISM.** * **Knowing HOW MANY is not knowing WHICH.** * **Round numbers matching is weak evidence — unless you read the constant.** * **My own last-turn result is a hypothesis too.** * **A global partition can understate a per-owner one.** * **A residual is measured against a population — name it.** ## Searching and tooling * **A search that returns thousands has no power; state the reach.** * **A substring match is not a hit.** **A regex miss looks like a null result — print one raw sample before believing a zero.** * **A derived table can be a cross product — measure its shape first.** * **After refuting an instrument, sweep everything that depended on it.** * **Before measuring how wrong a tool is, read what the tool actually does.** * **The instrument must pass its own control.** * **Classify a bulk before mining it. The residual is the prize.** * **Rank by similarity — the cliff is the finding.** But **read the values before trusting the rank.** * **Grep the nouns before designing the experiment — and believe it.** * **Grep gives you a file list — read *every* file on it.** * **The answer is often already in the doc that owns the subject — read it end to end.** A 🟡 often names its own route. * **Before re-trying a blocked idea, check whether the blocker's own doc already tried it.** * **Ship a regenerator with every artefact.** An artefact that moves by a pure reorder is a tool bug. * **Never print per-entry lines from a disc-wide sweep — aggregate.** ## Reading the data * **Read what a loader NAMES, not where it stores.** * **A field the disc never values still gets named by the loader.** * **An indexed read beats a deduped-pool adjacency read.** * **Check the whole string set, not the one matching word.** * **A dict keyed by record name across a multi-entry pak is a lie.** * **A set-difference over names hides reuse — join per USER.** * **A self-index names records, not files.** * **Case-insensitive hashing means two spellings can be one entry.** * **An "unresolved" name may be the wrong kind, namespace or prefix — or part of a cut asset.** * **A garbled value may be a real string in another encoding.** * **Two of my own counts disagreeing is a grammar clue.** * **A game's own typo is a join key.** * **A bias constant in the code is a join key.** * **A prefix trap: enumerate maximal `[A-Za-z0-9_]` runs, not `startswith`.** * **Re-deriving a format is not a finding — asking whether its values *resolve* is.** ## Mechanics that have bitten * **Never hand-convert a decimal VA — print `hex()`.** * **`grep -c` counts LINES** — use `grep -o | wc -l`. * **`Counter.most_common()` tie-breaks by insertion order — use `sorted()`.** * **Raw grep cannot see inside compressed pak entries.** * **Commit messages go in a file** (`git commit -F`); a literal `|` in a table cell needs escaping; `git log --all -- ` can hang. ## Runtime / emulator * **Look at the PNG** — and check its dimensions. * **"Animating" is not "still in a mission".** * **Dedup entity enumerations by position value.** * **Do not diagnose timing or liveness under gdb.** `ps %cpu` is cumulative. * **Classify screens by whole-image statistics, not named pixels** — a named pixel is only valid while the image sits at a known place, and nothing errors when it moves. * **Do not poll faster than the guest updates** — it manufactures a clean curve out of noise. * **A probe that never performs the action will "prove" the action does not exist.** * **The container's Canary binary can be older than the Canary source tree, and the failure mode is a hang, not an error.** After a merge into `sylpheed-re` the prebuilt `xenia_canary` had no `log_ui_draws`, no `mem_watch`, no `create_profile_if_none` — and an unknown cvar makes xenia open an SDL message box before logging is up, which headless is an unexplained freeze. Check before trusting a harness flag: `nm -C | grep cvars::`, and `build-canary Release` if it is missing. * **A capture armed *at* a screen only ever sees the steady state.** Anything about how a screen is built or animated has to be armed *before* it exists. Re-arming every few seconds and keeping every log tiles the approach: each F10 opens a new numbered file and closes the previous one complete. * **Measure animation in submitted frames, not in seconds.** `VdSwap` counts are the guest's own frames, so an emulator at 80 % of real time does not move them; a stopwatch reading does, silently and by an unknown factor. * **~~This game's menus drop d-pad presses shorter than ~0.3 s.~~ WITHDRAWN 2026-08-28 — the menu WRAPS, and I had not measured that.** The claim came from reading a cursor that ended up "one item short"; once wrap-around at both ends was measured ([`menu-navigation-semantics.md`](menu-navigation-semantics.md)), every one of those press counts is exactly right — four presses at 0.12 s moved four steps *through the bottom*, which lands one above where a non-wrapping menu would put it. **No press was ever dropped.** The real lesson is the general one: *a step count is only readable once you know the topology*, and I invented a hardware-flakiness story rather than testing the ends of the list. Still true and worth keeping: screenshot after every step and read the cursor, rather than trusting arithmetic over the press count. * **A screenshot taken right after a transition can catch a screen mid-fade.** A grab 2.5 s after Ⓑ returned the game to the title showed the title art with no `PRESS Ⓐ BUTTON` plate; one second later the plate was there. That very nearly went into the corpus as "the returned title has no plate". Sample a changing screen several times before writing down what it does *not* contain. * **Do not identify a menu cursor by label brightness.** The obvious oracle — "the focused label is the brightest row" — fails on this game's menus, because the background art is brighter behind some rows than the highlight is. It confidently named the wrong item on a frame whose ring was plainly elsewhere. Detect the **focus ring** in the gutter left of the labels instead (`tools/re-capture/menu_focus.py`, 254 vs <82 — no threshold tuning needed), and look at the PNG before believing either. * **`screenshot` samples at 0.5 Hz — it cannot time an animation.** Measured: ~2 s per grab (an `import` of the root plus an ImageMagick crop). A 0.4 s fade falls entirely between two samples, which is why a 40-frame burst across a screen change looked like an instant cut. For anything timed, record the display instead: `ffmpeg -f x11grab -framerate 30 -video_size x -i :98+, -t `, then read per-frame statistics off the file. Take the geometry from `xwininfo -root -tree`, the same way `bin/screenshot` does. * **A screen's brightness curve is not its fade quad.** The incoming screen's own elements animate in *after* the transition quad has cleared, so mean luminance keeps rising long after the fade is over — 1.47 s against a declared 0.97 s on one screen. Time the fade from where the frame is *pure black*, and take the ramp itself from the keyframes.