method: liveness has a second form -- failing correctly over a fraction of the corpus

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Wuu56cE8vJGTBtn1ppsk8v
This commit is contained in:
sylph-decoder
2026-08-31 03:33:00 +00:00
parent 976d657808
commit b6771fe9c0

View File

@@ -3096,3 +3096,25 @@ not wrong numbers**. My own worst instance — `ring_row.py`'s calibration, wron
because it was fitted against the wrong reference rows — would still not be caught, because it was fitted against the wrong reference rows — would still not be caught,
because nothing had retracted it: **nobody knew it was wrong.** This closes the because nothing had retracted it: **nobody knew it was wrong.** This closes the
class found in their tree and not the one found in mine. class found in their tree and not the one found in mine.
## Liveness has a second form: a checker that fails correctly over a fraction of the corpus
Every earlier instance in this corpus was a checker that **could not fail**.
`sylpheed-port` found the other shape: `audit-kinds` **fails correctly** and was
auditing **16 of 71** authored justifications, never saying so — while its clean
runs were being quoted as evidence the authored data is grounded. That was a
statement about a sixth of it.
📌 Their formulation is the one to keep: ***"I checked and it was fine"* and *"I
checked the part that declared itself"* read identically in a log.**
✅ Measured the same thing here. `check_refuted.py` covers **83 of 86** refutation-
shaped bullets — **97 %**, better than theirs and **equally unstated until now**.
It prints its scope before its verdict.
⚠️ **The remaining 3 are deliberate, and forcing them would be worse than the gap.**
They quote their claim in backticks and are bare identifiers (`+0x29d0`,
`position = instance 0x12c`); registering those would match every live mention of
the same offset. Both agents landed on the same rule independently: **report the
ratio, do not demand it be 1** — a counter that must be satisfied invites
mislabelling, which is a worse failure than an honest gap.