sylpheed-port found audit-kinds auditing 16 of 71 authored justifications and never saying so -- a checker that fails correctly while describing a sixth of the corpus. Their line is the one that generalises: 'I checked and it was fine' and 'I checked the part that declared itself' read identically in a log, and only one of them is what gets quoted. Measured here: of 86 refutation-shaped bullets in REFUTED.md, 83 are in the registered form. 97 %, which is better than their 16/71 but was equally unstated. The three gaps are deliberate, not a bug. They quote their claim in backticks and are bare identifiers -- +0x29d0, position = instance - 0x12c -- so registering them would match every live mention of the same offset and train the check to be ignored. Reported rather than forced to 100 %, for the same reason they report the ratio instead of demanding it: forcing a counter invites mislabelling, which is worse than the gap. Selftest and the real run both still exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Wuu56cE8vJGTBtn1ppsk8v
15 KiB
Executable File
15 KiB
Executable File