This repository has been archived on 2026-09-16. You can view files and clone it. You cannot open issues or pull requests or push a commit.
Files
Syplheed-Reborn/docs/re/METHOD.md
Sylpheed RE agent 88b3ce9af5 re: which GP_TITLE build is which screen, measured against the game
Q2. The archive is eight screens shipped twice, English and Japanese --
not the "build 4 title, 5 main menu, 6/8/9 submenus" the handoff claimed.
Build 8 is the JAPANESE main menu; 6 and 9 are the EN and JP EXTRAS, and
EXTRAS is the only submenu GP_TITLE holds. The PRESS (A) BUTTON plate is
its own build (2/3), composited over the title art and faded in a beat
later, not a state of build 4.

Confirmed by booting to the main menu and walking it: title, PRESS (A),
main menu and EXTRAS each match their render element for element. Builds
0/1 and 10/11 -- a DELTASABER / SYLPHEED A.I. plate -- were looked for in
the whole boot filmstrip, every title-side screen and the attract loop,
and appear in none of them; the reach of that negative is written down
rather than filled in with a guess.

Two rig traps went into METHOD: the menus drop d-pad presses shorter than
~0.3 s, and a grab 2.5 s after a transition can catch a screen mid-fade
-- which nearly wrote "the returned title has no plate" into the corpus.
2026-08-28 16:55:33 +00:00

133 lines
7.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Method traps already paid for
Each line cost an iteration at least once. They are general — they are not about
Sylpheed, they are about how this kind of measurement goes wrong.
Like [`REFUTED.md`](REFUTED.md), this list had been living in the autonomous
agent's loop prompt, i.e. nowhere durable. See [`README.md`](README.md) for the
✅/🟡/❔ confidence convention itself.
## Controls
* **Every result needs a control. A control that fails kills the instrument.**
* **Run the known-positive through a new filter FIRST.** Three filters have been
killed by their own control. When one fails, **read the known-good's
disassembly** before assuming a shape.
* **A measured negative is a result** — but a negative is only as strong as the
route you ran, so **state its reach**.
* **A null result needs its cause shown to have happened.**
* **A result with NO unknowns is suspicious.**
* **Census the whole set; always run the other population as the control.**
**Zero partials is stronger than a majority.**
* **A 2×2 partition is the sharpest general tool** — both off-diagonals empty is
a law.
* **Re-derive a doc's own numbers as the control.**
## Inference
* **Never conclude from ONE sample.**
* **A law proved on one population is a hypothesis on the next.**
* **Finding one exception does not imply a family.**
* **Consistency is not proof. A suggestive coincidence is a coincidence until
measured.** **An analogy is not a measurement.**
* **Same layout ≠ same instance.** **Same record-name set ≠ same object.**
* **A marker is only proven by what it leaves out.**
* **A high-confidence SCORE is not a high-confidence MECHANISM.**
* **Knowing HOW MANY is not knowing WHICH.**
* **Round numbers matching is weak evidence — unless you read the constant.**
* **My own last-turn result is a hypothesis too.**
* **A global partition can understate a per-owner one.**
* **A residual is measured against a population — name it.**
## Searching and tooling
* **A search that returns thousands has no power; state the reach.**
* **A substring match is not a hit.** **A regex miss looks like a null result —
print one raw sample before believing a zero.**
* **A derived table can be a cross product — measure its shape first.**
* **After refuting an instrument, sweep everything that depended on it.**
* **Before measuring how wrong a tool is, read what the tool actually does.**
* **The instrument must pass its own control.**
* **Classify a bulk before mining it. The residual is the prize.**
* **Rank by similarity — the cliff is the finding.** But **read the values
before trusting the rank.**
* **Grep the nouns before designing the experiment — and believe it.**
* **Grep gives you a file list — read *every* file on it.**
* **The answer is often already in the doc that owns the subject — read it end
to end.** A 🟡 often names its own route.
* **Before re-trying a blocked idea, check whether the blocker's own doc already
tried it.**
* **Ship a regenerator with every artefact.** An artefact that moves by a pure
reorder is a tool bug.
* **Never print per-entry lines from a disc-wide sweep — aggregate.**
## Reading the data
* **Read what a loader NAMES, not where it stores.**
* **A field the disc never values still gets named by the loader.**
* **An indexed read beats a deduped-pool adjacency read.**
* **Check the whole string set, not the one matching word.**
* **A dict keyed by record name across a multi-entry pak is a lie.**
* **A set-difference over names hides reuse — join per USER.**
* **A self-index names records, not files.**
* **Case-insensitive hashing means two spellings can be one entry.**
* **An "unresolved" name may be the wrong kind, namespace or prefix — or part of
a cut asset.**
* **A garbled value may be a real string in another encoding.**
* **Two of my own counts disagreeing is a grammar clue.**
* **A game's own typo is a join key.**
* **A bias constant in the code is a join key.**
* **A prefix trap: enumerate maximal `[A-Za-z0-9_]` runs, not `startswith`.**
* **Re-deriving a format is not a finding — asking whether its values *resolve*
is.**
## Mechanics that have bitten
* **Never hand-convert a decimal VA — print `hex()`.**
* **`grep -c` counts LINES** — use `grep -o | wc -l`.
* **`Counter.most_common()` tie-breaks by insertion order — use `sorted()`.**
* **Raw grep cannot see inside compressed pak entries.**
* **Commit messages go in a file** (`git commit -F`); a literal `|` in a table
cell needs escaping; `git log --all -- <path>` can hang.
## Runtime / emulator
* **Look at the PNG** — and check its dimensions.
* **"Animating" is not "still in a mission".**
* **Dedup entity enumerations by position value.**
* **Do not diagnose timing or liveness under gdb.** `ps %cpu` is cumulative.
* **Classify screens by whole-image statistics, not named pixels** — a named
pixel is only valid while the image sits at a known place, and nothing errors
when it moves.
* **Do not poll faster than the guest updates** — it manufactures a clean curve
out of noise.
* **A probe that never performs the action will "prove" the action does not
exist.**
* **The container's Canary binary can be older than the Canary source tree, and
the failure mode is a hang, not an error.** After a merge into `sylpheed-re`
the prebuilt `xenia_canary` had no `log_ui_draws`, no `mem_watch`, no
`create_profile_if_none` — and an unknown cvar makes xenia open an SDL message
box before logging is up, which headless is an unexplained freeze. Check before
trusting a harness flag: `nm -C <binary> | grep cvars::<flag>`, and
`build-canary Release` if it is missing.
* **A capture armed *at* a screen only ever sees the steady state.** Anything
about how a screen is built or animated has to be armed *before* it exists.
Re-arming every few seconds and keeping every log tiles the approach: each F10
opens a new numbered file and closes the previous one complete.
* **Measure animation in submitted frames, not in seconds.** `VdSwap` counts are
the guest's own frames, so an emulator at 80 % of real time does not move them;
a stopwatch reading does, silently and by an unknown factor.
* **This game's menus drop d-pad presses shorter than ~0.3 s.** `pad.py dpad`
defaults to 0.06 s, and its docstring says longer "auto-repeats and
overshoots". On the title main menu that default is *dropped*: four presses at
0.12 s moved the cursor one step, four at 0.20 s moved it none, while
0.30/0.50/0.80 s each moved it exactly one step and none of them repeated. A
scripted navigation that comes out one item short is this, not a wrong item
count — screenshot after every step and check the cursor rather than trusting
the press count.
* **A screenshot taken right after a transition can catch a screen mid-fade.**
A grab 2.5 s after Ⓑ returned the game to the title showed the title art with
no `PRESS Ⓐ BUTTON` plate; one second later the plate was there. That very
nearly went into the corpus as "the returned title has no plate". Sample a
changing screen several times before writing down what it does *not* contain.