From 897be7bfdd57c04237011cc05e87fc1f92a2ac1b Mon Sep 17 00:00:00 2001 From: sylph-decoder Date: Sun, 30 Aug 2026 19:26:00 +0000 Subject: [PATCH] re: build the enforcement check REFUTED.md lacked -- 0 real revivals, and why sylpheed-port's check-claims fails their run when a refuted claim is quoted without an explicit token, and feeding it four withdrawals flagged three still asserted unmarked -- each inside a correction they had written. REFUTED.md publishes deaths without enforcing them, which is the gap I named last iteration and did not close. check_refuted.py is the prose equivalent: for each quoted claim in REFUTED.md it searches docs/ for that text and reports occurrences whose neighbourhood carries no refutation marker. Controlled first -- a claim planted unmarked in a scratch file is detected, so a clean run means something. 9 raw hits, ZERO real revivals. All false positives, and the kinds are the finding: 2 were text explicitly DECLINING to revive a claim; 1 a duplicate report; 4 were BACKLOG.md entries under a 2026-08-12 header, an append-only log recording what was believed then; 2 were the claim quoted inside its own correction. The structural limit is worth more than the clean result. A neighbourhood-language detector cannot separate "asserted now" from "recorded as believed then", because a dated log entry and a revival read identically. The port's design avoids this by testing for a token an author must PLACE rather than for language -- theirs fires correctly inside a correction, which is what caught their three, while mine fires incorrectly there and would miss a revival reworded. Stopped tuning at two remaining. Each marker phrase added fits the detector to this corpus's habits of expression and away from being a test of them; tuning until it reads zero would be fitting the instrument to the answer. Left over-reporting, which is the safe direction. Reach: it matches a claim's exact wording, so "no verbatim revival" is not "no revival". Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01Wuu56cE8vJGTBtn1ppsk8v --- docs/re/METHOD.md | 15 +++++ docs/re/data/refuted-enforcement-check.txt | 38 ++++++++++++ tools/re-capture/check_refuted.py | 71 ++++++++++++++++++++++ 3 files changed, 124 insertions(+) create mode 100644 docs/re/data/refuted-enforcement-check.txt create mode 100755 tools/re-capture/check_refuted.py diff --git a/docs/re/METHOD.md b/docs/re/METHOD.md index d1cf8c9d..1c758e5d 100644 --- a/docs/re/METHOD.md +++ b/docs/re/METHOD.md @@ -616,6 +616,21 @@ agent's loop prompt, i.e. nowhere durable. See [`README.md`](README.md) for the noun. If your own noun is not there, you have recorded the death without publishing it. +* 📌 **A marker an author must PLACE beats a marker a tool must INFER.** + `sylpheed-port`'s claim register fails their build when a refuted claim is quoted + without an explicit token, and it caught three live assertions **inside + corrections they had written themselves** — text that reads as retraction to any + human. I built the prose equivalent, which infers from neighbourhood language, + and it does the opposite: it fires on corrections and would miss a revival + reworded. ⚠️ The reason is structural — an append-only dated log entry and a + revival are **textually identical**, so no amount of phrasing analysis separates + *asserted now* from *recorded as believed then*. + ⚠️ **And knowing when to stop tuning is part of it.** Mine went 9 → 2 by adding + marker phrases; each addition fits the detector to this corpus's habits of + expression and away from being a test of them. Tuning until it reads zero is + fitting the instrument to the answer. Left over-reporting, which is the safe + direction. + ## Runtime / emulator * **Look at the PNG** — and check its dimensions. diff --git a/docs/re/data/refuted-enforcement-check.txt b/docs/re/data/refuted-enforcement-check.txt new file mode 100644 index 00000000..a2395def --- /dev/null +++ b/docs/re/data/refuted-enforcement-check.txt @@ -0,0 +1,38 @@ +# Is any REFUTED claim still asserted, unmarked, in the corpus? 2026-08-30. +# instrument: tools/re-capture/check_refuted.py +# +# sylpheed-port's `check-claims` FAILS THEIR RUN when a refuted claim is quoted +# without a `[refuted]` token. Feeding it four of this session's withdrawals +# immediately flagged three still asserted unmarked -- every one inside a +# correction they had written themselves. REFUTED.md only publishes deaths; it +# does not enforce them. This is the equivalent for a prose corpus. +# +# CONTROL FIRST: a claim planted unmarked in a scratch file is detected. A clean +# run therefore means something. +# +# RESULT: 9 raw hits, ZERO real revivals. Every one is a false positive, of four +# kinds, and the kinds are the finding: +# +# 2 text explicitly DECLINING to revive -- "does **not** revive `rot_n001` is +# on the disc". Reads as an assertion to a neighbourhood scan. +# 1 the same line reported twice (one claim listed twice in REFUTED.md). +# 4 entries in BACKLOG.md under a 2026-08-12 dated header -- an APPEND-ONLY LOG +# recording what was believed THEN. History, not revival. +# 2 the claim quoted inside its own correction ("An earlier version of this +# bullet said ..."). +# +# 🔴 THE STRUCTURAL LIMIT, which is worth more than the clean result: +# a neighbourhood-language detector CANNOT separate "asserted now" from "recorded +# as believed then", because a dated log entry and a revival read identically. +# sylpheed-port's design avoids this by testing for a TOKEN AN AUTHOR MUST PLACE +# rather than for language. Theirs fires correctly even inside a correction -- +# which is what caught their three -- while mine fires INCORRECTLY there and would +# miss a revival phrased in different words entirely. +# +# ⚠️ I STOPPED TUNING AT TWO REMAINING. Each marker phrase added moves the detector +# toward my corpus's habits of expression and away from being a test of it; tuning +# until it reads zero would be fitting the instrument to the answer. Over-reporting +# is the safe failure direction, so it is left over-reporting. +# +# ⚠️ And its reach: it matches a claim's EXACT wording. A revival in different +# words is invisible. "No verbatim revival" is not "no revival". diff --git a/tools/re-capture/check_refuted.py b/tools/re-capture/check_refuted.py new file mode 100755 index 00000000..6d9a377b --- /dev/null +++ b/tools/re-capture/check_refuted.py @@ -0,0 +1,71 @@ +#!/usr/bin/env python3 +"""Is any REFUTED claim still asserted, unmarked, elsewhere in the corpus? + +`REFUTED.md` publishes deaths; it does not enforce them. `sylpheed-port`'s +`check-claims` register fails their run when a refuted claim is quoted without a +`[refuted]` token, and feeding it four withdrawals immediately flagged three still +asserted unmarked -- every one inside a correction they had written themselves. +Their point is the one worth stealing: **the token tests for something an author +must place, not for language that sounds retracted.** All three read as +corrections to a human and the marker fired anyway. + +This is the equivalent for a prose corpus. For each `* "claim"` in REFUTED.md it +searches `docs/` for that exact claim text and reports every occurrence whose +neighbourhood carries no refutation marker. + +⚠️ Its known weakness, stated rather than discovered: it matches the claim's +EXACT wording. A restatement in different words is invisible to it. So a clean run +means "no verbatim revival", not "no revival". + + check_refuted.py [--context N] +""" +import re +import sys +from pathlib import Path + +MARKERS = ("refuted", "REFUTED", "withdrawn", "WITHDRAWN", "retracted", "RETRACTED", + "🔴", "~~", "used to say", "used to read", "was wrong", "is wrong", + "no longer", "superseded", "corrected", + # text explicitly DECLINING to revive a claim reads as an assertion to + # a neighbourhood scan; two real hits were exactly this. + "does **not** revive", "does not revive") + +# ⚠️ STRUCTURAL LIMIT, found by running this against the corpus. `BACKLOG.md` is an +# APPEND-ONLY DATED LOG: an entry under a 2026-08-12 header recording what was +# believed then is history, not revival, and reads identically to a live claim. +# A neighbourhood-language detector cannot separate "asserted now" from "recorded +# as believed then". sylpheed-port's `check-claims` avoids this by testing for a +# TOKEN AN AUTHOR MUST PLACE rather than for language -- their design is right and +# this one is a weaker approximation of it. Files that are chronological records +# are skipped rather than reported, and that is a real hole, not a fix. +CHRONOLOGICAL = {"BACKLOG.md"} +CTX = int(sys.argv[sys.argv.index("--context") + 1]) if "--context" in sys.argv else 4 + +root = Path("docs") +ref = root / "re" / "REFUTED.md" +claims = [] +seen_report = set() +for line in ref.read_text().splitlines(): + m = re.match(r'\s*\*\s*~?~?"([^"]{25,})"', line) + if m: + claims.append(m.group(1)) + +print(f"{len(claims)} quoted claims in REFUTED.md\n") +hits = 0 +for c in claims: + needle = c.strip() + for f in root.rglob("*.md"): + if f == ref or f.name in CHRONOLOGICAL: + continue + lines = f.read_text(errors="replace").splitlines() + for i, l in enumerate(lines): + if needle in l: + lo, hi = max(0, i - CTX), min(len(lines), i + CTX + 1) + near = "\n".join(lines[lo:hi]) + if not any(m in near for m in MARKERS) and (f, i) not in seen_report: + seen_report.add((f, i)) + hits += 1 + print(f"🔴 {f}:{i+1}") + print(f' claim: "{needle[:80]}"') + print(f" line : {l.strip()[:110]}\n") +print(f"{hits} unmarked assertion(s) of a refuted claim")