port: the register check had no executable control, and an empty register passed forever

check-claims guards the refuted register, the thing both agents lean on when they
say a dead claim is not being re-asserted, and it had no control machinery at all.
Every 'planted a revival, it failed, removed it, it passed' in DECISIONS was done
by hand, once, and never again -- in a repository where two of my own tools carry
the line 'a control that does not execute is not a control'. I wrote that about
somebody else's tool.

The hole the Decoder found in their equivalent was here too. The scan loop runs
once per register row; with no rows it runs zero times, fail stays 0, and the
script printed 'every refuted claim appears only inside its correction' and exited
0. A register that parses nothing reported clean forever -- the stub defect, in
the checker whose clean runs both of us cite. It now exits 2 with 'the harness is
broken, not the corpus'.

--control executes four cases, each driving this script as a subprocess and
reading its real exit code: clean 0, unmarked revival 1, marked revival 0 with no
false positive, empty register 2. Asserting in check-all.

Two things taken from their build of the same thing rather than invented: the
self-test drives the real machinery and reads its actual exit code -- my first
--selftest reasoned about what the harness would do, which is the cheaper mistake
and the one I made -- and the three-way exit convention, which is what lets 'the
corpus is dirty' and 'the checker is broken' be different answers instead of both
being nonzero.

The plant lands in a real scanned directory, because a control that runs somewhere
the tool does not look proves nothing about the tool. Verified two-directionally:
pointing the plant at an unscanned path makes the control report itself broken.

Still without harness self-tests and filed rather than left looking finished:
audit-kinds and verify-transcode-fidelity.

Every asserting check passes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF
This commit is contained in:
Sylpheed port agent
2026-08-31 01:22:43 +00:00
parent 75be660fb4
commit 4fff1beecd
4 changed files with 113 additions and 2 deletions

View File

@@ -151,7 +151,7 @@ HANDOFF.
| Milestone | Needs | HANDOFF | State |
|---|---|---|---|
| all — control harnesses that assert themselves | **nothing from anybody; three tools still lack it** | `d38adcf` | 🟡 **DONE FOR `contract-check`, NOT for the rest.** `--selftest` feeds the machinery a stub that cannot fail and requires it to be flagged; exit codes separate **0** all good / **1** a real check failed / **2** the harness is broken. Asserting in `check-all`. ⚠️ `check-claims`, `audit-kinds` and `verify-transcode-fidelity` have controls and **no harness self-test** — the shape is known and the fix is cheap, and this row exists so the gap does not read as finished. 🔴 The self-test caught two defects while being written: a first version that *argued* the harness would flag the stub instead of measuring it, and a `src` selection that anchored anything outside one list at the wrong document, flagging the stub for a fabricated reason. |
| all — control harnesses that assert themselves | **nothing from anybody; TWO tools still lack it** | `d38adcf` | 🟡 **DONE FOR `contract-check` AND `check-claims`, not for the rest.** `check-claims --control` now executes four cases as subprocesses — clean 0, unmarked revival 1, marked revival 0, **empty register 2** — where before it had **no control machinery at all** and an empty register reported clean forever. Verified two-directionally: pointing the plant at an unscanned path makes the control report itself broken. Remaining: `audit-kinds`, `verify-transcode-fidelity`. Earlier text: 🟡 **DONE FOR `contract-check`, NOT for the rest.** `--selftest` feeds the machinery a stub that cannot fail and requires it to be flagged; exit codes separate **0** all good / **1** a real check failed / **2** the harness is broken. Asserting in `check-all`. ⚠️ `check-claims`, `audit-kinds` and `verify-transcode-fidelity` have controls and **no harness self-test** — the shape is known and the fix is cheap, and this row exists so the gap does not read as finished. 🔴 The self-test caught two defects while being written: a first version that *argued* the harness would flag the stub instead of measuring it, and a `src` selection that anchored anything outside one list at the wrong document, flagging the stub for a fabricated reason. |
## Coverage hole in my own check, 2026-08-31 — derived from HANDOFF `0159527`

View File

@@ -9,7 +9,7 @@ dies, which is what this file is for.
<!-- INDEX: generated by tools/port/index-decisions -- do not hand-edit -->
266 sections. Search this before re-deriving anything.
267 sections. Search this before re-deriving anything.
* [P0 — the exporter, 2026-08-28](#p0--the-exporter-2026-08-28)
* [P1 — Godot draws the screen, 2026-08-28](#p1--godot-draws-the-screen-2026-08-28)
@@ -277,6 +277,7 @@ dies, which is what this file is for.
* [🔴 RETRACTED: the `S00A` coverage hole was my control's filter, not the check](#retracted-the-s00a-coverage-hole-was-my-controls-filter-not-the-check)
* [Their two tools had the shape I shipped, and the general form is sharper now](#their-two-tools-had-the-shape-i-shipped-and-the-general-form-is-sharper-now)
* [Closing the two-directional gap: the control harness now asserts itself](#closing-the-two-directional-gap-the-control-harness-now-asserts-itself)
* [The register check had no executable control, and an empty register passed forever](#the-register-check-had-no-executable-control-and-an-empty-register-passed-forever)
<!-- /INDEX -->
## P0 — the exporter, 2026-08-28
@@ -13640,3 +13641,45 @@ screens.
`verify-transcode-fidelity` have controls but no harness self-test. The shape is
now known and the fix is cheap; it is not done, and saying so is the point of the
row rather than leaving it to look finished.
## The register check had no executable control, and an empty register passed forever
`check-claims` guards the refuted register — the thing both agents lean on when
they say a dead claim is not being re-asserted. It had **no control machinery at
all**. Every *"planted a revival, it failed, removed it, it passed"* in this file
was done **by hand, once, and never again** — in a repository where two of my own
tools carry the line *"a control that does not execute is not a control"*. I wrote
that about somebody else's tool.
### 🔴 And the hole the Decoder found in theirs was here too
The scan loop runs once per register row. **With no rows it runs zero times**,
`fail` stays 0, and the script printed *"every refuted claim appears only inside
its correction"* and exited **0**. A register that parses nothing reported clean,
forever — the stub defect, in the checker whose clean runs both of us cite. It now
exits **2** with *"the harness is broken, not the corpus"*.
### Four cases, executed, driving the real script as a subprocess
| case | exit |
|---|---|
| clean tree | **0** |
| unmarked revival planted | **1** |
| revival planted **marked** | **0** — and no false positive |
| register emptied | **2** |
📌 Two things taken from the Decoder's build of the same thing rather than
invented: **the self-test drives the real machinery and reads its actual exit
code** — my first `--selftest` reasoned about what the harness *would* do, which
is the cheaper mistake and the one I made — and **the three-way exit convention**,
which is what lets "the corpus is dirty" and "the checker is broken" be different
answers instead of both being "nonzero".
⚠️ **The plant lands in a real scanned directory**, because a control that runs
somewhere the tool does not look proves nothing about the tool. Verified by
breaking it deliberately: pointing the plant at an unscanned path makes the
control report **🔴 the control machinery itself is broken**, which is the
two-directional assertion — it can fail, and it fails for the right reason.
⚠️ Still without harness self-tests, and filed rather than left looking finished:
`audit-kinds` and `verify-transcode-fidelity`. Same shape, cheap, not done.

View File

@@ -65,6 +65,11 @@ step menu-audio must-pass env OUT="$OUT/audio" tools/port/verify-menu-audi
step decisions-index must-pass tools/port/index-decisions --check
# A refuted claim asserted outside its correction is a lie the corpus tells a
# reader who greps for it. Registered claims must carry an explicit `[refuted]`.
# 🔴 The register check had NO executable control until 2026-08-31 -- every
# "planted a revival and it failed" in DECISIONS was done by hand, once. Four
# cases now drive it as a subprocess and read its real exit code, including an
# EMPTY REGISTER, which used to report clean forever.
step claims-control must-pass tools/port/check-claims --control
step refuted-claims must-pass tools/port/check-claims
echo
echo "reported, not asserted:"

View File

@@ -58,6 +58,69 @@ HANDOFF Q10 says nothing on the disc
ROWS
)
[ -n "${CLAIMS_REGISTER+x}" ] && REGISTER="$CLAIMS_REGISTER"
# 🔴 A REGISTER THAT PARSES NOTHING REPORTED CLEAN, FOREVER. The scan loop runs
# once per row; with no rows it runs zero times, `fail` stays 0, and the script
# printed "every refuted claim appears only inside its correction" and exited 0.
# That is the stub defect -- prints a result, asserts nothing -- sitting in the
# checker whose clean runs both agents lean on. The Decoder found it in their
# equivalent the same day; it was here too.
_rows=$(printf '%s\n' "$REGISTER" | grep -c '[^[:space:]]' || true)
if [ "$_rows" -eq 0 ]; then
echo "🔴 the refuted register is EMPTY -- this check would pass everything." >&2
echo " Exit 2: the harness is broken, not the corpus." >&2
exit 2
fi
# --------------------------------------------------------------------------
# `--control`: the known negatives, EXECUTED.
#
# 🔴 Until now this check had NO control machinery at all. Every "planted a
# revival, it failed, removed it, it passed" in `DECISIONS.md` was done BY HAND,
# once, and never again -- in a repository where two of my own tools carry the
# line *"a control that does not execute is not a control"*. It was written
# about somebody else's tool.
#
# Four cases, each driving THIS script as a subprocess and reading its real exit
# code rather than reasoning about what it would do:
#
# clean tree -> 0
# unmarked revival planted -> 1 (the check must catch it)
# revival planted MARKED -> 0 (and must not false-positive on it)
# register emptied -> 2 (the harness is broken, not the corpus)
#
# The plant lands in a real scanned directory, because a control that runs
# somewhere the tool does not look proves nothing about the tool.
if [ "${1:-}" = "--control" ]; then
probe="docs/port/.claims-control-probe.md"
trap 'rm -f "$probe"' EXIT INT TERM
claim=$(printf '%s\n' "$REGISTER" | grep -m1 '[^[:space:]]')
ok=0
run() { CLAIMS_CONTROL=1 "$0" >/dev/null 2>&1; echo $?; }
rm -f "$probe"
for case in "clean::0" "unmarked:$claim:1" "marked:$claim [refuted]:0"; do
IFS=: read -r name body want <<<"$case"
if [ -n "$body" ]; then printf '%s\n' "$body" > "$probe"; else rm -f "$probe"; fi
got=$(run)
if [ "$got" = "$want" ]; then
printf ' %-26s exit %s ✅\n' "$name" "$got"
else
printf ' %-26s exit %s, wanted %s 🔴\n' "$name" "$got" "$want"; ok=1
fi
done
rm -f "$probe"
got=$(CLAIMS_REGISTER="" "$0" >/dev/null 2>&1; echo $?)
if [ "$got" = "2" ]; then printf ' %-26s exit 2 ✅\n' "empty register"
else printf ' %-26s exit %s, wanted 2 🔴\n' "empty register" "$got"; ok=1; fi
echo
[ $ok -eq 0 ] && echo "the register check fails when it must, and says so distinctly" \
|| echo "🔴 the control machinery itself is broken"
exit $ok
fi
# ─── THE WITHDRAWAL-TIME HOOK ────────────────────────────────────────────────
# The register enforces claims it KNOWS ABOUT; knowing about them was manual, and
# that is how ~8 claims were withdrawn this session and 0 registered. A sweep