re: build a shared probe harness so the same lessons stop being re-learned

Four probes were written from a blank file and each re-learned the same lessons
by losing a run: that a flat run cannot be told from a frozen guest without a
stall witness, that results held to the end of a run are destroyed by a turn
timeout, that a roster count which is not the stage's member count means a
different stage loaded and must be discarded, and that a run's witness state has
to be read before its numbers. Writing each lesson down did not stop the next
probe repeating it, because each probe started from nothing.

probeharness.py makes them structural. Probe(baseline=N) discovers the roster,
rescans up to five times and refuses to start if the count never reaches the
baseline. The witness is calibrated on construction, sampled by tick() and
reported by status() and summary(), so a probe cannot forget it, and when no
witness is found it reports UNVALIDATED rather than zero stalls. emit() flushes
on every line. craft(), strengths(), alive() and heap() supply the
roster-to-craft link, per-record liveness and the raw heap, so a new probe
writes only its own logic.

Verified rather than asserted: deploy_probe.py reimplements the per-record
deployment watch on top of it in about forty lines against wave7_probe's
hundred and fifty, and its first live run was clean -- 116 roster records, 32
witnesses at 10/s, zero stalled samples, seven losses tracked, and the TSV
written incrementally. Nothing about the result is new, which is the point: the
harness reproduces a known-good measurement.

The existing probes are deliberately not ported. They work, and rewriting them
would risk changing results other documents cite. New probes should use the
harness; old ones should be ported when they next need a change.
This commit is contained in:
Sylpheed RE agent
2026-08-25 05:22:24 +00:00
parent 738df50803
commit feb535a8fb
5 changed files with 347 additions and 0 deletions

View File

@@ -630,6 +630,18 @@ search cannot find a *schedule*.
that cannot be told from a freeze**; the recurring fix is the shared harness
noted earlier. Needs two HUD readings at *different* values in non-stalled
samples.
***(2026-08-25) SHARED PROBE HARNESS built and verified**
([`probe-harness.md`](probe-harness.md), `tools/re-capture/probeharness.py`).
Makes structural the four lessons that were re-learned in four separate probes:
**built-in stall witness** (says `UNVALIDATED` when absent rather than reporting
zero stalls), **`emit()` flushes every line** so a timeout cannot destroy
results, **baseline discard with rescan** (`Probe(baseline=116)` refuses to
start on a different stage), and `summary()` putting the witness first.
Also provides `craft()`/`strengths()`/`alive()`/`heap()` so a new probe writes
only its own logic. ✅ Verified: `deploy_probe.py` reimplements the deployment
watch in ~40 lines vs 150, first live run clean — 116 roster, 32 witnesses at
10/s, **0 stalled**, 7 losses, TSV written incrementally. ⚠️ Existing probes
deliberately **not** ported — they work and other docs cite their results.
* ~~🚧 BLOCKER: t=210/240 unreachable in one turn~~ — **superseded, see above**;
it rested on an untested assumption that a turn is one shell call. 595 s shell cap ~220 s boot (a ~190 s title movie that cannot be
tapped through) ~25 s startup = **~350 s observation ≈ 193 game-seconds**.