From 575e287526c3b8abb64a73377d88bb900e0b4fa2 Mon Sep 17 00:00:00 2001 From: Sylpheed port agent Date: Mon, 31 Aug 2026 01:22:43 +0000 Subject: [PATCH] port: the register check had no executable control, and an empty register passed forever check-claims guards the refuted register, the thing both agents lean on when they say a dead claim is not being re-asserted, and it had no control machinery at all. Every 'planted a revival, it failed, removed it, it passed' in DECISIONS was done by hand, once, and never again -- in a repository where two of my own tools carry the line 'a control that does not execute is not a control'. I wrote that about somebody else's tool. The hole the Decoder found in their equivalent was here too. The scan loop runs once per register row; with no rows it runs zero times, fail stays 0, and the script printed 'every refuted claim appears only inside its correction' and exited 0. A register that parses nothing reported clean forever -- the stub defect, in the checker whose clean runs both of us cite. It now exits 2 with 'the harness is broken, not the corpus'. --control executes four cases, each driving this script as a subprocess and reading its real exit code: clean 0, unmarked revival 1, marked revival 0 with no false positive, empty register 2. Asserting in check-all. Two things taken from their build of the same thing rather than invented: the self-test drives the real machinery and reads its actual exit code -- my first --selftest reasoned about what the harness would do, which is the cheaper mistake and the one I made -- and the three-way exit convention, which is what lets 'the corpus is dirty' and 'the checker is broken' be different answers instead of both being nonzero. The plant lands in a real scanned directory, because a control that runs somewhere the tool does not look proves nothing about the tool. Verified two-directionally: pointing the plant at an unscanned path makes the control report itself broken. Still without harness self-tests and filed rather than left looking finished: audit-kinds and verify-transcode-fidelity. Every asserting check passes. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01N7FiFFFwbvG2uxdcEh8HyF --- docs/port/BLOCKED.md | 2 +- docs/port/DECISIONS.md | 45 ++++++++++++++++++++++++++++- tools/port/check-all | 5 ++++ tools/port/check-claims | 63 +++++++++++++++++++++++++++++++++++++++++ 4 files changed, 113 insertions(+), 2 deletions(-) diff --git a/docs/port/BLOCKED.md b/docs/port/BLOCKED.md index 2876e5b5..732a8988 100644 --- a/docs/port/BLOCKED.md +++ b/docs/port/BLOCKED.md @@ -151,7 +151,7 @@ HANDOFF. | Milestone | Needs | HANDOFF | State | |---|---|---|---| -| all — control harnesses that assert themselves | **nothing from anybody; three tools still lack it** | `d38adcf` | 🟡 **DONE FOR `contract-check`, NOT for the rest.** `--selftest` feeds the machinery a stub that cannot fail and requires it to be flagged; exit codes separate **0** all good / **1** a real check failed / **2** the harness is broken. Asserting in `check-all`. ⚠️ `check-claims`, `audit-kinds` and `verify-transcode-fidelity` have controls and **no harness self-test** — the shape is known and the fix is cheap, and this row exists so the gap does not read as finished. 🔴 The self-test caught two defects while being written: a first version that *argued* the harness would flag the stub instead of measuring it, and a `src` selection that anchored anything outside one list at the wrong document, flagging the stub for a fabricated reason. | +| all — control harnesses that assert themselves | **nothing from anybody; TWO tools still lack it** | `d38adcf` | 🟡 **DONE FOR `contract-check` AND `check-claims`, not for the rest.** `check-claims --control` now executes four cases as subprocesses — clean 0, unmarked revival 1, marked revival 0, **empty register 2** — where before it had **no control machinery at all** and an empty register reported clean forever. Verified two-directionally: pointing the plant at an unscanned path makes the control report itself broken. Remaining: `audit-kinds`, `verify-transcode-fidelity`. Earlier text: 🟡 **DONE FOR `contract-check`, NOT for the rest.** `--selftest` feeds the machinery a stub that cannot fail and requires it to be flagged; exit codes separate **0** all good / **1** a real check failed / **2** the harness is broken. Asserting in `check-all`. ⚠️ `check-claims`, `audit-kinds` and `verify-transcode-fidelity` have controls and **no harness self-test** — the shape is known and the fix is cheap, and this row exists so the gap does not read as finished. 🔴 The self-test caught two defects while being written: a first version that *argued* the harness would flag the stub instead of measuring it, and a `src` selection that anchored anything outside one list at the wrong document, flagging the stub for a fabricated reason. | ## Coverage hole in my own check, 2026-08-31 — derived from HANDOFF `0159527` diff --git a/docs/port/DECISIONS.md b/docs/port/DECISIONS.md index 7c40b1b9..ea4a48d9 100644 --- a/docs/port/DECISIONS.md +++ b/docs/port/DECISIONS.md @@ -9,7 +9,7 @@ dies, which is what this file is for. -266 sections. Search this before re-deriving anything. +267 sections. Search this before re-deriving anything. * [P0 — the exporter, 2026-08-28](#p0--the-exporter-2026-08-28) * [P1 — Godot draws the screen, 2026-08-28](#p1--godot-draws-the-screen-2026-08-28) @@ -277,6 +277,7 @@ dies, which is what this file is for. * [🔴 RETRACTED: the `S00A` coverage hole was my control's filter, not the check](#retracted-the-s00a-coverage-hole-was-my-controls-filter-not-the-check) * [Their two tools had the shape I shipped, and the general form is sharper now](#their-two-tools-had-the-shape-i-shipped-and-the-general-form-is-sharper-now) * [Closing the two-directional gap: the control harness now asserts itself](#closing-the-two-directional-gap-the-control-harness-now-asserts-itself) +* [The register check had no executable control, and an empty register passed forever](#the-register-check-had-no-executable-control-and-an-empty-register-passed-forever) ## P0 — the exporter, 2026-08-28 @@ -13640,3 +13641,45 @@ screens. `verify-transcode-fidelity` have controls but no harness self-test. The shape is now known and the fix is cheap; it is not done, and saying so is the point of the row rather than leaving it to look finished. + +## The register check had no executable control, and an empty register passed forever + +`check-claims` guards the refuted register — the thing both agents lean on when +they say a dead claim is not being re-asserted. It had **no control machinery at +all**. Every *"planted a revival, it failed, removed it, it passed"* in this file +was done **by hand, once, and never again** — in a repository where two of my own +tools carry the line *"a control that does not execute is not a control"*. I wrote +that about somebody else's tool. + +### 🔴 And the hole the Decoder found in theirs was here too + +The scan loop runs once per register row. **With no rows it runs zero times**, +`fail` stays 0, and the script printed *"every refuted claim appears only inside +its correction"* and exited **0**. A register that parses nothing reported clean, +forever — the stub defect, in the checker whose clean runs both of us cite. It now +exits **2** with *"the harness is broken, not the corpus"*. + +### Four cases, executed, driving the real script as a subprocess + +| case | exit | +|---|---| +| clean tree | **0** | +| unmarked revival planted | **1** | +| revival planted **marked** | **0** — and no false positive | +| register emptied | **2** | + +📌 Two things taken from the Decoder's build of the same thing rather than +invented: **the self-test drives the real machinery and reads its actual exit +code** — my first `--selftest` reasoned about what the harness *would* do, which +is the cheaper mistake and the one I made — and **the three-way exit convention**, +which is what lets "the corpus is dirty" and "the checker is broken" be different +answers instead of both being "nonzero". + +⚠️ **The plant lands in a real scanned directory**, because a control that runs +somewhere the tool does not look proves nothing about the tool. Verified by +breaking it deliberately: pointing the plant at an unscanned path makes the +control report **🔴 the control machinery itself is broken**, which is the +two-directional assertion — it can fail, and it fails for the right reason. + +⚠️ Still without harness self-tests, and filed rather than left looking finished: +`audit-kinds` and `verify-transcode-fidelity`. Same shape, cheap, not done. diff --git a/tools/port/check-all b/tools/port/check-all index 6f055654..d34915b8 100755 --- a/tools/port/check-all +++ b/tools/port/check-all @@ -65,6 +65,11 @@ step menu-audio must-pass env OUT="$OUT/audio" tools/port/verify-menu-audi step decisions-index must-pass tools/port/index-decisions --check # A refuted claim asserted outside its correction is a lie the corpus tells a # reader who greps for it. Registered claims must carry an explicit `[refuted]`. +# 🔴 The register check had NO executable control until 2026-08-31 -- every +# "planted a revival and it failed" in DECISIONS was done by hand, once. Four +# cases now drive it as a subprocess and read its real exit code, including an +# EMPTY REGISTER, which used to report clean forever. +step claims-control must-pass tools/port/check-claims --control step refuted-claims must-pass tools/port/check-claims echo echo "reported, not asserted:" diff --git a/tools/port/check-claims b/tools/port/check-claims index 962a7440..fe54d6eb 100755 --- a/tools/port/check-claims +++ b/tools/port/check-claims @@ -58,6 +58,69 @@ HANDOFF Q10 says nothing on the disc ROWS ) +[ -n "${CLAIMS_REGISTER+x}" ] && REGISTER="$CLAIMS_REGISTER" + +# 🔴 A REGISTER THAT PARSES NOTHING REPORTED CLEAN, FOREVER. The scan loop runs +# once per row; with no rows it runs zero times, `fail` stays 0, and the script +# printed "every refuted claim appears only inside its correction" and exited 0. +# That is the stub defect -- prints a result, asserts nothing -- sitting in the +# checker whose clean runs both agents lean on. The Decoder found it in their +# equivalent the same day; it was here too. +_rows=$(printf '%s\n' "$REGISTER" | grep -c '[^[:space:]]' || true) +if [ "$_rows" -eq 0 ]; then + echo "🔴 the refuted register is EMPTY -- this check would pass everything." >&2 + echo " Exit 2: the harness is broken, not the corpus." >&2 + exit 2 +fi + +# -------------------------------------------------------------------------- +# `--control`: the known negatives, EXECUTED. +# +# 🔴 Until now this check had NO control machinery at all. Every "planted a +# revival, it failed, removed it, it passed" in `DECISIONS.md` was done BY HAND, +# once, and never again -- in a repository where two of my own tools carry the +# line *"a control that does not execute is not a control"*. It was written +# about somebody else's tool. +# +# Four cases, each driving THIS script as a subprocess and reading its real exit +# code rather than reasoning about what it would do: +# +# clean tree -> 0 +# unmarked revival planted -> 1 (the check must catch it) +# revival planted MARKED -> 0 (and must not false-positive on it) +# register emptied -> 2 (the harness is broken, not the corpus) +# +# The plant lands in a real scanned directory, because a control that runs +# somewhere the tool does not look proves nothing about the tool. +if [ "${1:-}" = "--control" ]; then + probe="docs/port/.claims-control-probe.md" + trap 'rm -f "$probe"' EXIT INT TERM + claim=$(printf '%s\n' "$REGISTER" | grep -m1 '[^[:space:]]') + ok=0 + run() { CLAIMS_CONTROL=1 "$0" >/dev/null 2>&1; echo $?; } + rm -f "$probe" + for case in "clean::0" "unmarked:$claim:1" "marked:$claim [refuted]:0"; do + IFS=: read -r name body want <<<"$case" + if [ -n "$body" ]; then printf '%s\n' "$body" > "$probe"; else rm -f "$probe"; fi + got=$(run) + if [ "$got" = "$want" ]; then + printf ' %-26s exit %s ✅\n' "$name" "$got" + else + printf ' %-26s exit %s, wanted %s 🔴\n' "$name" "$got" "$want"; ok=1 + fi + done + rm -f "$probe" + got=$(CLAIMS_REGISTER="" "$0" >/dev/null 2>&1; echo $?) + if [ "$got" = "2" ]; then printf ' %-26s exit 2 ✅\n' "empty register" + else printf ' %-26s exit %s, wanted 2 🔴\n' "empty register" "$got"; ok=1; fi + echo + [ $ok -eq 0 ] && echo "the register check fails when it must, and says so distinctly" \ + || echo "🔴 the control machinery itself is broken" + exit $ok +fi + + + # ─── THE WITHDRAWAL-TIME HOOK ──────────────────────────────────────────────── # The register enforces claims it KNOWS ABOUT; knowing about them was manual, and # that is how ~8 claims were withdrawn this session and 0 registered. A sweep