re: F1 dynamic attempt -- harness debugged through four bugs, still no number
Some checks failed
CI / Native — linux (pull_request) Failing after 35m32s
CI / WASM — Web (pull_request) Successful in 31m31s
CI / Formatting (pull_request) Failing after 1m29s

Tried to close out issue #1 with an actual draw-log measurement this
iteration: built f1_hold_capture.py to boot to the settled main menu via the
proven glyph-gated route, arm the F10 UI-draw capture, hold a direction, and
read cursor position per frame -- the instrument f1-menu-repeat-harness-
built-not-answered.md already validated but never got to run against a
reachable menu.

Four bugs found across four boot attempts:

1. tap() shelled out to pad.py without this script's own env, so the press
   went to /tmp/xenia_pad.txt while Canary watched OUT/pad.txt -- an
   unobserved press indistinguishable from a dead pad. Fixed with an
   in-process tap() using the same pad() the hold uses; confirmed working
   the next run (title 154.5s, menu 163.0s).

2. ui_draw_capture_frames/max were persisted at 3/20000 from a prior
   session in xenia-canary.config.toml -- log_ui_draws is now a documented
   no-op (F10 arms unconditionally) and these two cvars didn't visibly
   respond to command-line overrides. Bumped to 600/400000 directly in the
   config.

3. The real blocker: this container has no signed-in profile (no content/
   directory at all -- a fresh container after a restart, which every
   container is right after one). Without a profile the title's sign-in
   dialog sets IsUIActive() true, which reproduces
   structures/title-a-press-fault.md's already-diagnosed unbounded-
   keystroke-queue crash -- verified byte-for-byte against that page's own
   addresses (PC 0x868 past sub_82457038, 0x828F3xxx registers, identical
   host/guest address arithmetic), looping continuously from before F10 was
   ever pressed. Fixed by creating a profile (--create_profile_if_none) and
   signing in (--logged_profile_slot_0_xuid), matching boot_menu.sh, which
   already did this and so never hit it. Confirmed: zero crashes with the
   fix, dozens per run without it.

4. Found but not re-verified: no xsetroot blank before launch, so a stale
   X-root frame from a killed prior run gave a false "TITLE" read at 2.6s,
   before any real window existed -- skip_intro.sh already blanks the root
   for exactly this reason. Fixed in the script.

Ran out of budget before a clean end-to-end run landed. Still no number for
issue #1 -- the Port keeps -1.0. Flagged prominently (HANDOFF, REFUTED.md)
because bug 3 will hit any bare run-canary invocation in any fresh
container, not just this script.
This commit is contained in:
sylph-decoder
2026-09-12 12:02:53 +00:00
parent dfbefff512
commit 089aa8d7c2
4 changed files with 326 additions and 0 deletions

View File

@@ -6593,3 +6593,17 @@ the coarse screen-diff detector undercounted" reading over the
Keystroke/400-100ms one, though neither is confirmed yet and the actual
producer still isn't found. Still no number — still don't take one from
here.
## 🔴 F1 dynamic attempt this iteration — harness debugged through four bugs, still no number
[`docs/re/f1-hold-capture-harness-debugged.md`](../re/f1-hold-capture-harness-debugged.md).
Tried to actually run the measurement this time. Found and fixed three real
bugs (an env mismatch that silently dropped a scripted press, a stale
persisted capture-window cvar, and — the big one — a fresh container with no
signed-in profile reproducing the already-documented
`structures/title-a-press-fault.md` crash in a loop) and found-but-didn't-
re-verify a fourth (a stale X-root frame giving a false-positive screen
read). Ran out of iteration budget before a clean end-to-end run. **Still no
number for you** — keep `-1.0`. The harness should work next attempt;
flagging for whoever's container hits the same profile-crash first, since
every fresh container starts with none.

View File

@@ -1094,3 +1094,16 @@ attached to a right direction; only the reason is refuted here.
emulator needed — or re-run with a repeat-capable file driver and read
cursor position off the draw log per frame, not a coarse screen-diff.
[`f1-no-repeat-was-the-harness.md`](f1-no-repeat-was-the-harness.md) ⟨harness⟩
* "F10-arming a UI draw capture is safe on any settled screen" — **implicit
in every prior use of it, and false in a fresh container.** A `run-canary`
launch with no signed-in profile (every container right after a restart —
no `content/` directory exists yet) reproduces
[`structures/title-a-press-fault.md`](structures/title-a-press-fault.md)'s
already-diagnosed `IsUIActive`/unbounded-keystroke-queue crash in a tight
loop, starting before F10 is ever pressed — it is the profile, not the
hotkey. Reproduced byte-for-byte against that page's own addresses (crash
PC `0x868` past `sub_82457038`, registers in the `0x828F3xxx` input-manager
range, `0x1701D0000``0x701D0000` host/guest arithmetic identical).
`boot_menu.sh` already signs in a profile; a bare `run-canary` call, like
`f1_hold_capture.py`'s first drafts, does not, and hits this every time in
a fresh container. [`f1-hold-capture-harness-debugged.md`](f1-hold-capture-harness-debugged.md) ⟨environment⟩

View File

@@ -0,0 +1,107 @@
# F1 dynamic attempt — harness debugged through four bugs, no measurement yet
**Status:** ❔ not answered. Built
[`tools/re-capture/f1_hold_capture.py`](../../tools/re-capture/f1_hold_capture.py)
to close out issue #1 with an actual draw-log measurement — boot to the
settled main menu (the proven glyph-gated route), arm the F10 UI-draw
capture, hold a direction, read the cursor's position per frame. Never
completed a clean end-to-end run this iteration; four real bugs found along
the way, three fixed and confirmed, one fixed but not re-verified. Same shape
as [`f1-menu-repeat-harness-built-not-answered.md`](f1-menu-repeat-harness-built-not-answered.md) —
the harness is closer, the number still isn't here.
⚠️ **Container reset warning, read this if you are a fresh session:** bug 3
below will hit *any* scripted `run-canary` invocation in a container that has
never signed in a profile — which is every container right after a restart,
per this project's own "your session is new each time" note. `boot_menu.sh`
already handles it; a bare `run-canary` call does not.
## Bug 1 — `tap()` shelled out to `pad.py` with the wrong environment
First attempt's "tap A" used `subprocess.run([sys.executable, PAD, ...])`
without passing this script's own `env` dict, so the subprocess inherited
`os.environ` instead — where `XENIA_PAD_FILE` was never set at the OS level,
only in the local dict. `pad.py` wrote its default `/tmp/xenia_pad.txt`;
Canary was watching `OUT/pad.txt`. The press never arrived: **an
unobserved press that looks identical to a dead pad**, and it cost one full
520 s boot before the mismatch was found (only one `[file-pad]` log line the
whole run — the initial clear).
**Fixed:** an in-process `tap()` using the same `pad()` function the hold
uses, no subprocess. **Confirmed working** the very next run: title reached
at 154.5 s, tap delivered, menu reached at 163.0 s.
## Bug 2 — the persisted `ui_draw_capture_frames` was 3, disc-wide default
`ui_draw_capture_frames = 3` / `ui_draw_capture_max = 20000` sat in
`~/.local/share/Xenia/xenia-canary.config.toml` from a previous session — an
old capture's window, kept because `--log_ui_draws=true` (the flag most
existing scripts pass) is **now a documented no-op** (`command_processor.cc`:
*"OBSOLETE — the UI draw-order capture is armed by F10 unconditionally now"*).
Passing `--ui_draw_capture_frames=3000` on the command line did not visibly
change the value logged at startup, and this container's saved config wins
for these two cvars regardless of the flag.
**Fixed:** edited the toml directly, `600` / `400000` — wide enough to cover
a multi-second hold at any plausible present rate. Confirmed applied (log
dump shows the new values) on the next run, though it wasn't the run that
also isolated bug 3, so it hasn't independently produced a capture yet.
## Bug 3 — no signed-in profile, and it reproduces the documented crash
The real blocker. This container had **no `content/` directory at all**
a fresh container after a restart, exactly as the loop brief warns, and
`f1_hold_capture.py` (unlike `boot_menu.sh`) never checked. Without a
profile, the title shows a sign-in dialog, `IsUIActive()` goes true, and —
**this is the exact crash `structures/title-a-press-fault.md` already
diagnosed** — `XamInputGetKeystrokeEx` returns success with an empty
keystroke forever, the game's unbounded pump queues every empty result, and
it eventually crashes on a bad allocation.
Reproduced, not just matched by description: the crash dump's PC
(`0x824578A0`) sits **0x868 past `sub_82457038`** (that doc's "keystroke
pump"), registers hold values in the `0x828F3xxx` range (the input-manager
singleton at `0x828F3888`), and the access violation address relationship —
`0x1701D0000` (host) → `0x701D0000` (guest, `r9`) — matches that doc's own
worked arithmetic exactly. Without a profile the crash repeated in a tight
loop (one dump roughly every 200 log lines) for the entire run, starting
before F10 was ever pressed — this is not an F10 side effect, it is the
title's own boot path with no profile to sign into.
**Fixed:** `run-canary --create_profile_if_none=F1Probe`, wait ~5 s, kill it
— creates `content/<XUID>/`. Script now resolves that XUID and passes
`--logged_profile_slot_0_xuid=<XUID>`, matching `boot_menu.sh`. **Confirmed**:
zero crash dumps in the run that used it, against dozens per run before.
## Bug 4 — no root-window blank before launch; found, not re-verified
The run after fixing bug 3 detected "TITLE" at **2.6 s** — long before any
real window could exist — and never recovered. `skip_intro.sh`'s own header
explains why: *"The X root keeps the DEAD session's last frame, so a fresh
launch would be detected as already at the title."* The previous run's
killed window left stale pixels on the shared `:98` display; my script never
blanked it, unlike every proven boot script.
**Fixed in the script** (`xsetroot -solid black` before each launch) but
**not re-run this iteration** — four bugs and roughly nine minutes of boot
time across four attempts is where this stopped rather than chasing a fifth.
## What's left
The harness should now be correct: env-safe input, a wide enough capture
window, a signed-in profile, a blanked root. The next run is the actual
measurement — hold a direction on the settled menu, read `read_draws.py`'s
per-frame quad positions, group by `(page, width, height)` shape (the
approach `menu_repeat_probe.sh`'s own notes call "track one quad shape's
position over time," the one that correctly reports no motion when there is
none), and report which shape-group's position changes, at what frame
spacing, at the achieved present rate.
## Reach
Everything here is ⟨harness⟩ / ⟨environment⟩ — about this container and this
tool, not about the game. Bug 3's crash mechanism is not new (already fully
diagnosed in `structures/title-a-press-fault.md`); what's new is the trigger
(a freshly-reset container with no profile) and that it reproduces
byte-for-byte against that page's own addresses. No claim about F1's actual
repeat rate is made or changed here.