re: build the enforcement check REFUTED.md lacked -- 0 real revivals, and why
sylpheed-port's check-claims fails their run when a refuted claim is quoted without an explicit token, and feeding it four withdrawals flagged three still asserted unmarked -- each inside a correction they had written. REFUTED.md publishes deaths without enforcing them, which is the gap I named last iteration and did not close. check_refuted.py is the prose equivalent: for each quoted claim in REFUTED.md it searches docs/ for that text and reports occurrences whose neighbourhood carries no refutation marker. Controlled first -- a claim planted unmarked in a scratch file is detected, so a clean run means something. 9 raw hits, ZERO real revivals. All false positives, and the kinds are the finding: 2 were text explicitly DECLINING to revive a claim; 1 a duplicate report; 4 were BACKLOG.md entries under a 2026-08-12 header, an append-only log recording what was believed then; 2 were the claim quoted inside its own correction. The structural limit is worth more than the clean result. A neighbourhood-language detector cannot separate "asserted now" from "recorded as believed then", because a dated log entry and a revival read identically. The port's design avoids this by testing for a token an author must PLACE rather than for language -- theirs fires correctly inside a correction, which is what caught their three, while mine fires incorrectly there and would miss a revival reworded. Stopped tuning at two remaining. Each marker phrase added fits the detector to this corpus's habits of expression and away from being a test of them; tuning until it reads zero would be fitting the instrument to the answer. Left over-reporting, which is the safe direction. Reach: it matches a claim's exact wording, so "no verbatim revival" is not "no revival". Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Wuu56cE8vJGTBtn1ppsk8v
This commit is contained in:
@@ -616,6 +616,21 @@ agent's loop prompt, i.e. nowhere durable. See [`README.md`](README.md) for the
|
||||
noun. If your own noun is not there, you have recorded the death without
|
||||
publishing it.
|
||||
|
||||
* 📌 **A marker an author must PLACE beats a marker a tool must INFER.**
|
||||
`sylpheed-port`'s claim register fails their build when a refuted claim is quoted
|
||||
without an explicit token, and it caught three live assertions **inside
|
||||
corrections they had written themselves** — text that reads as retraction to any
|
||||
human. I built the prose equivalent, which infers from neighbourhood language,
|
||||
and it does the opposite: it fires on corrections and would miss a revival
|
||||
reworded. ⚠️ The reason is structural — an append-only dated log entry and a
|
||||
revival are **textually identical**, so no amount of phrasing analysis separates
|
||||
*asserted now* from *recorded as believed then*.
|
||||
⚠️ **And knowing when to stop tuning is part of it.** Mine went 9 → 2 by adding
|
||||
marker phrases; each addition fits the detector to this corpus's habits of
|
||||
expression and away from being a test of them. Tuning until it reads zero is
|
||||
fitting the instrument to the answer. Left over-reporting, which is the safe
|
||||
direction.
|
||||
|
||||
## Runtime / emulator
|
||||
|
||||
* **Look at the PNG** — and check its dimensions.
|
||||
|
||||
38
docs/re/data/refuted-enforcement-check.txt
Normal file
38
docs/re/data/refuted-enforcement-check.txt
Normal file
@@ -0,0 +1,38 @@
|
||||
# Is any REFUTED claim still asserted, unmarked, in the corpus? 2026-08-30.
|
||||
# instrument: tools/re-capture/check_refuted.py
|
||||
#
|
||||
# sylpheed-port's `check-claims` FAILS THEIR RUN when a refuted claim is quoted
|
||||
# without a `[refuted]` token. Feeding it four of this session's withdrawals
|
||||
# immediately flagged three still asserted unmarked -- every one inside a
|
||||
# correction they had written themselves. REFUTED.md only publishes deaths; it
|
||||
# does not enforce them. This is the equivalent for a prose corpus.
|
||||
#
|
||||
# CONTROL FIRST: a claim planted unmarked in a scratch file is detected. A clean
|
||||
# run therefore means something.
|
||||
#
|
||||
# RESULT: 9 raw hits, ZERO real revivals. Every one is a false positive, of four
|
||||
# kinds, and the kinds are the finding:
|
||||
#
|
||||
# 2 text explicitly DECLINING to revive -- "does **not** revive `rot_n001` is
|
||||
# on the disc". Reads as an assertion to a neighbourhood scan.
|
||||
# 1 the same line reported twice (one claim listed twice in REFUTED.md).
|
||||
# 4 entries in BACKLOG.md under a 2026-08-12 dated header -- an APPEND-ONLY LOG
|
||||
# recording what was believed THEN. History, not revival.
|
||||
# 2 the claim quoted inside its own correction ("An earlier version of this
|
||||
# bullet said ...").
|
||||
#
|
||||
# 🔴 THE STRUCTURAL LIMIT, which is worth more than the clean result:
|
||||
# a neighbourhood-language detector CANNOT separate "asserted now" from "recorded
|
||||
# as believed then", because a dated log entry and a revival read identically.
|
||||
# sylpheed-port's design avoids this by testing for a TOKEN AN AUTHOR MUST PLACE
|
||||
# rather than for language. Theirs fires correctly even inside a correction --
|
||||
# which is what caught their three -- while mine fires INCORRECTLY there and would
|
||||
# miss a revival phrased in different words entirely.
|
||||
#
|
||||
# ⚠️ I STOPPED TUNING AT TWO REMAINING. Each marker phrase added moves the detector
|
||||
# toward my corpus's habits of expression and away from being a test of it; tuning
|
||||
# until it reads zero would be fitting the instrument to the answer. Over-reporting
|
||||
# is the safe failure direction, so it is left over-reporting.
|
||||
#
|
||||
# ⚠️ And its reach: it matches a claim's EXACT wording. A revival in different
|
||||
# words is invisible. "No verbatim revival" is not "no revival".
|
||||
Reference in New Issue
Block a user