f6-out-of-sample-RESULT: three of six predictions failed — shape the residue #9

Open
opened 2026-09-04 15:48:08 +00:00 by fabi · 1 comment
Owner

docs/re/f6-out-of-sample-RESULT.md opens:

"three of six predictions failed""the check I built last iteration fails its first independent test."

That is live, it is on a branch nobody has reviewed, and it was not in the original phase-6 list.

This item is a shaping request, not the work. The Decoder should say what remains open in its own words, what instrument would close it, and whether the failing check is worth repairing or withdrawing — then propose items for whatever survives. A check that fails its first independent test is the more urgent half: an instrument that reports pass when the feature is absent is the failure mode this project has now met three times.

Pass: either concrete follow-on items with stated instruments, or a written argument that nothing here needs doing — both are answers; silence is not.

`docs/re/f6-out-of-sample-RESULT.md` opens: > **"three of six predictions failed"** … **"the check I built last iteration fails its first independent test."** That is live, it is on a branch nobody has reviewed, and it was not in the original phase-6 list. **This item is a shaping request, not the work.** The Decoder should say what remains open in its own words, what instrument would close it, and whether the failing check is worth repairing or withdrawing — then propose items for whatever survives. A check that fails its first independent test is the more urgent half: an instrument that reports pass when the feature is absent is the failure mode this project has now met three times. **Pass:** either concrete follow-on items with stated instruments, or a written argument that nothing here needs doing — both are answers; silence is not.
fabi added this to the Title screen milestone 2026-09-04 15:48:08 +00:00
fabi added the agent/decoderstate/proposedkind/item labels 2026-09-04 15:48:08 +00:00
fabi added state/approved and removed state/proposed labels 2026-09-04 15:59:19 +00:00
sylph-decoder added state/needs-human and removed state/approved labels 2026-09-11 21:26:29 +00:00
Collaborator

Answered: docs/re/f6-residue-shaping.md,
in #23 (322869b).

In my own words: the residue isn't three independent misses. Two of the three
failures (the ptcopyright/parent ramp ratio, the sweep-leads-plate lead)
move together — same direction, 0.75× and 0.71× of predicted — and
f6-unit10 already decoded that the sweep and plate families are gated by
separately triggered parents, so nothing guarantees their relative phase
is a constant across boots the way a ratio within one family is (which is
exactly the three that held). The gross mislabeling behind the original
1.7× conflict was fixed two days before the prereg, so it's ruled out here.
The third failure (the amplitude curve fit) is a different, single-element
self-consistency check and isn't explained by the same argument.

What would close it: more independent no-input boots (filed as #25,
state/proposed — new dynamic RE, not this pass's own static scope) to tell
"genuine boot-to-boot phase variance" from "instrument artifact" apart —
clustering supports the former, scatter tracking pulse_period's own noise
supports the latter.

check_labels.py: repaired, not patched. Split into gating identity checks
(3/3 including f6c) and informational timing checks (0/2) rather than one
verdict conflating "wrong element" with "timing isn't constant." Tolerances
untouched. Verified against synthetic data shaped like the real residue, and
confirmed the selftest's injected mislabel still fails an identity check —
no capture survives in this container to run it against real logs.

Nothing here reinstates the withdrawn 0.138–0.141 lead. Moving to
state/needs-human; I can't close my own work.

Answered: [`docs/re/f6-residue-shaping.md`](../../docs/re/f6-residue-shaping.md), in #23 (322869b). In my own words: the residue isn't three independent misses. Two of the three failures (the `ptcopyright`/parent ramp ratio, the sweep-leads-plate lead) move together — same direction, 0.75× and 0.71× of predicted — and `f6-unit10` already decoded that the sweep and plate families are gated by **separately triggered** parents, so nothing guarantees their relative phase is a constant across boots the way a ratio *within* one family is (which is exactly the three that held). The gross mislabeling behind the *original* 1.7× conflict was fixed two days before the prereg, so it's ruled out here. The third failure (the amplitude curve fit) is a different, single-element self-consistency check and isn't explained by the same argument. What would close it: more independent no-input boots (filed as #25, `state/proposed` — new dynamic RE, not this pass's own static scope) to tell "genuine boot-to-boot phase variance" from "instrument artifact" apart — clustering supports the former, scatter tracking `pulse_period`'s own noise supports the latter. `check_labels.py`: repaired, not patched. Split into gating identity checks (3/3 including `f6c`) and informational timing checks (0/2) rather than one verdict conflating "wrong element" with "timing isn't constant." Tolerances untouched. Verified against synthetic data shaped like the real residue, and confirmed the selftest's injected mislabel still fails an identity check — no capture survives in this container to run it against real logs. Nothing here reinstates the withdrawn 0.138–0.141 lead. Moving to `state/needs-human`; I can't close my own work.
Sign in to join this conversation.
2 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: fabi/Sylpheed#9