port: the register check had no executable control, and an empty register passed forever
check-claims guards the refuted register, the thing both agents lean on when they say a dead claim is not being re-asserted, and it had no control machinery at all. Every 'planted a revival, it failed, removed it, it passed' in DECISIONS was done by hand, once, and never again -- in a repository where two of my own tools carry the line 'a control that does not execute is not a control'. I wrote that about somebody else's tool. The hole the Decoder found in their equivalent was here too. The scan loop runs once per register row; with no rows it runs zero times, fail stays 0, and the script printed 'every refuted claim appears only inside its correction' and exited 0. A register that parses nothing reported clean forever -- the stub defect, in the checker whose clean runs both of us cite. It now exits 2 with 'the harness is broken, not the corpus'. --control executes four cases, each driving this script as a subprocess and reading its real exit code: clean 0, unmarked revival 1, marked revival 0 with no false positive, empty register 2. Asserting in check-all. Two things taken from their build of the same thing rather than invented: the self-test drives the real machinery and reads its actual exit code -- my first --selftest reasoned about what the harness would do, which is the cheaper mistake and the one I made -- and the three-way exit convention, which is what lets 'the corpus is dirty' and 'the checker is broken' be different answers instead of both being nonzero. The plant lands in a real scanned directory, because a control that runs somewhere the tool does not look proves nothing about the tool. Verified two-directionally: pointing the plant at an unscanned path makes the control report itself broken. Still without harness self-tests and filed rather than left looking finished: audit-kinds and verify-transcode-fidelity. Every asserting check passes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF
This commit is contained in:
@@ -151,7 +151,7 @@ HANDOFF.
|
||||
|
||||
| Milestone | Needs | HANDOFF | State |
|
||||
|---|---|---|---|
|
||||
| all — control harnesses that assert themselves | **nothing from anybody; three tools still lack it** | `d38adcf` | 🟡 **DONE FOR `contract-check`, NOT for the rest.** `--selftest` feeds the machinery a stub that cannot fail and requires it to be flagged; exit codes separate **0** all good / **1** a real check failed / **2** the harness is broken. Asserting in `check-all`. ⚠️ `check-claims`, `audit-kinds` and `verify-transcode-fidelity` have controls and **no harness self-test** — the shape is known and the fix is cheap, and this row exists so the gap does not read as finished. 🔴 The self-test caught two defects while being written: a first version that *argued* the harness would flag the stub instead of measuring it, and a `src` selection that anchored anything outside one list at the wrong document, flagging the stub for a fabricated reason. |
|
||||
| all — control harnesses that assert themselves | **nothing from anybody; TWO tools still lack it** | `d38adcf` | 🟡 **DONE FOR `contract-check` AND `check-claims`, not for the rest.** `check-claims --control` now executes four cases as subprocesses — clean 0, unmarked revival 1, marked revival 0, **empty register 2** — where before it had **no control machinery at all** and an empty register reported clean forever. Verified two-directionally: pointing the plant at an unscanned path makes the control report itself broken. Remaining: `audit-kinds`, `verify-transcode-fidelity`. Earlier text: 🟡 **DONE FOR `contract-check`, NOT for the rest.** `--selftest` feeds the machinery a stub that cannot fail and requires it to be flagged; exit codes separate **0** all good / **1** a real check failed / **2** the harness is broken. Asserting in `check-all`. ⚠️ `check-claims`, `audit-kinds` and `verify-transcode-fidelity` have controls and **no harness self-test** — the shape is known and the fix is cheap, and this row exists so the gap does not read as finished. 🔴 The self-test caught two defects while being written: a first version that *argued* the harness would flag the stub instead of measuring it, and a `src` selection that anchored anything outside one list at the wrong document, flagging the stub for a fabricated reason. |
|
||||
|
||||
## Coverage hole in my own check, 2026-08-31 — derived from HANDOFF `0159527`
|
||||
|
||||
|
||||
@@ -9,7 +9,7 @@ dies, which is what this file is for.
|
||||
|
||||
<!-- INDEX: generated by tools/port/index-decisions -- do not hand-edit -->
|
||||
|
||||
266 sections. Search this before re-deriving anything.
|
||||
267 sections. Search this before re-deriving anything.
|
||||
|
||||
* [P0 — the exporter, 2026-08-28](#p0--the-exporter-2026-08-28)
|
||||
* [P1 — Godot draws the screen, 2026-08-28](#p1--godot-draws-the-screen-2026-08-28)
|
||||
@@ -277,6 +277,7 @@ dies, which is what this file is for.
|
||||
* [🔴 RETRACTED: the `S00A` coverage hole was my control's filter, not the check](#retracted-the-s00a-coverage-hole-was-my-controls-filter-not-the-check)
|
||||
* [Their two tools had the shape I shipped, and the general form is sharper now](#their-two-tools-had-the-shape-i-shipped-and-the-general-form-is-sharper-now)
|
||||
* [Closing the two-directional gap: the control harness now asserts itself](#closing-the-two-directional-gap-the-control-harness-now-asserts-itself)
|
||||
* [The register check had no executable control, and an empty register passed forever](#the-register-check-had-no-executable-control-and-an-empty-register-passed-forever)
|
||||
|
||||
<!-- /INDEX -->
|
||||
## P0 — the exporter, 2026-08-28
|
||||
@@ -13640,3 +13641,45 @@ screens.
|
||||
`verify-transcode-fidelity` have controls but no harness self-test. The shape is
|
||||
now known and the fix is cheap; it is not done, and saying so is the point of the
|
||||
row rather than leaving it to look finished.
|
||||
|
||||
## The register check had no executable control, and an empty register passed forever
|
||||
|
||||
`check-claims` guards the refuted register — the thing both agents lean on when
|
||||
they say a dead claim is not being re-asserted. It had **no control machinery at
|
||||
all**. Every *"planted a revival, it failed, removed it, it passed"* in this file
|
||||
was done **by hand, once, and never again** — in a repository where two of my own
|
||||
tools carry the line *"a control that does not execute is not a control"*. I wrote
|
||||
that about somebody else's tool.
|
||||
|
||||
### 🔴 And the hole the Decoder found in theirs was here too
|
||||
|
||||
The scan loop runs once per register row. **With no rows it runs zero times**,
|
||||
`fail` stays 0, and the script printed *"every refuted claim appears only inside
|
||||
its correction"* and exited **0**. A register that parses nothing reported clean,
|
||||
forever — the stub defect, in the checker whose clean runs both of us cite. It now
|
||||
exits **2** with *"the harness is broken, not the corpus"*.
|
||||
|
||||
### Four cases, executed, driving the real script as a subprocess
|
||||
|
||||
| case | exit |
|
||||
|---|---|
|
||||
| clean tree | **0** |
|
||||
| unmarked revival planted | **1** |
|
||||
| revival planted **marked** | **0** — and no false positive |
|
||||
| register emptied | **2** |
|
||||
|
||||
📌 Two things taken from the Decoder's build of the same thing rather than
|
||||
invented: **the self-test drives the real machinery and reads its actual exit
|
||||
code** — my first `--selftest` reasoned about what the harness *would* do, which
|
||||
is the cheaper mistake and the one I made — and **the three-way exit convention**,
|
||||
which is what lets "the corpus is dirty" and "the checker is broken" be different
|
||||
answers instead of both being "nonzero".
|
||||
|
||||
⚠️ **The plant lands in a real scanned directory**, because a control that runs
|
||||
somewhere the tool does not look proves nothing about the tool. Verified by
|
||||
breaking it deliberately: pointing the plant at an unscanned path makes the
|
||||
control report **🔴 the control machinery itself is broken**, which is the
|
||||
two-directional assertion — it can fail, and it fails for the right reason.
|
||||
|
||||
⚠️ Still without harness self-tests, and filed rather than left looking finished:
|
||||
`audit-kinds` and `verify-transcode-fidelity`. Same shape, cheap, not done.
|
||||
|
||||
@@ -65,6 +65,11 @@ step menu-audio must-pass env OUT="$OUT/audio" tools/port/verify-menu-audi
|
||||
step decisions-index must-pass tools/port/index-decisions --check
|
||||
# A refuted claim asserted outside its correction is a lie the corpus tells a
|
||||
# reader who greps for it. Registered claims must carry an explicit `[refuted]`.
|
||||
# 🔴 The register check had NO executable control until 2026-08-31 -- every
|
||||
# "planted a revival and it failed" in DECISIONS was done by hand, once. Four
|
||||
# cases now drive it as a subprocess and read its real exit code, including an
|
||||
# EMPTY REGISTER, which used to report clean forever.
|
||||
step claims-control must-pass tools/port/check-claims --control
|
||||
step refuted-claims must-pass tools/port/check-claims
|
||||
echo
|
||||
echo "reported, not asserted:"
|
||||
|
||||
@@ -58,6 +58,69 @@ HANDOFF Q10 says nothing on the disc
|
||||
ROWS
|
||||
)
|
||||
|
||||
[ -n "${CLAIMS_REGISTER+x}" ] && REGISTER="$CLAIMS_REGISTER"
|
||||
|
||||
# 🔴 A REGISTER THAT PARSES NOTHING REPORTED CLEAN, FOREVER. The scan loop runs
|
||||
# once per row; with no rows it runs zero times, `fail` stays 0, and the script
|
||||
# printed "every refuted claim appears only inside its correction" and exited 0.
|
||||
# That is the stub defect -- prints a result, asserts nothing -- sitting in the
|
||||
# checker whose clean runs both agents lean on. The Decoder found it in their
|
||||
# equivalent the same day; it was here too.
|
||||
_rows=$(printf '%s\n' "$REGISTER" | grep -c '[^[:space:]]' || true)
|
||||
if [ "$_rows" -eq 0 ]; then
|
||||
echo "🔴 the refuted register is EMPTY -- this check would pass everything." >&2
|
||||
echo " Exit 2: the harness is broken, not the corpus." >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
# --------------------------------------------------------------------------
|
||||
# `--control`: the known negatives, EXECUTED.
|
||||
#
|
||||
# 🔴 Until now this check had NO control machinery at all. Every "planted a
|
||||
# revival, it failed, removed it, it passed" in `DECISIONS.md` was done BY HAND,
|
||||
# once, and never again -- in a repository where two of my own tools carry the
|
||||
# line *"a control that does not execute is not a control"*. It was written
|
||||
# about somebody else's tool.
|
||||
#
|
||||
# Four cases, each driving THIS script as a subprocess and reading its real exit
|
||||
# code rather than reasoning about what it would do:
|
||||
#
|
||||
# clean tree -> 0
|
||||
# unmarked revival planted -> 1 (the check must catch it)
|
||||
# revival planted MARKED -> 0 (and must not false-positive on it)
|
||||
# register emptied -> 2 (the harness is broken, not the corpus)
|
||||
#
|
||||
# The plant lands in a real scanned directory, because a control that runs
|
||||
# somewhere the tool does not look proves nothing about the tool.
|
||||
if [ "${1:-}" = "--control" ]; then
|
||||
probe="docs/port/.claims-control-probe.md"
|
||||
trap 'rm -f "$probe"' EXIT INT TERM
|
||||
claim=$(printf '%s\n' "$REGISTER" | grep -m1 '[^[:space:]]')
|
||||
ok=0
|
||||
run() { CLAIMS_CONTROL=1 "$0" >/dev/null 2>&1; echo $?; }
|
||||
rm -f "$probe"
|
||||
for case in "clean::0" "unmarked:$claim:1" "marked:$claim [refuted]:0"; do
|
||||
IFS=: read -r name body want <<<"$case"
|
||||
if [ -n "$body" ]; then printf '%s\n' "$body" > "$probe"; else rm -f "$probe"; fi
|
||||
got=$(run)
|
||||
if [ "$got" = "$want" ]; then
|
||||
printf ' %-26s exit %s ✅\n' "$name" "$got"
|
||||
else
|
||||
printf ' %-26s exit %s, wanted %s 🔴\n' "$name" "$got" "$want"; ok=1
|
||||
fi
|
||||
done
|
||||
rm -f "$probe"
|
||||
got=$(CLAIMS_REGISTER="" "$0" >/dev/null 2>&1; echo $?)
|
||||
if [ "$got" = "2" ]; then printf ' %-26s exit 2 ✅\n' "empty register"
|
||||
else printf ' %-26s exit %s, wanted 2 🔴\n' "empty register" "$got"; ok=1; fi
|
||||
echo
|
||||
[ $ok -eq 0 ] && echo "the register check fails when it must, and says so distinctly" \
|
||||
|| echo "🔴 the control machinery itself is broken"
|
||||
exit $ok
|
||||
fi
|
||||
|
||||
|
||||
|
||||
# ─── THE WITHDRAWAL-TIME HOOK ────────────────────────────────────────────────
|
||||
# The register enforces claims it KNOWS ABOUT; knowing about them was manual, and
|
||||
# that is how ~8 claims were withdrawn this session and 0 registered. A sweep
|
||||
|
||||
Reference in New Issue
Block a user