# A shared probe harness — so the same lessons stop being re-learned Status: ✅ built and verified on a live run. ## Why Four separate probes were written from a blank file, and each re-learned the same lessons by losing a run: | lesson | where it was learned | where it was re-learned | |---|---|---| | a flat run is indistinguishable from a frozen guest without a **stall witness** | `guest-stalls.md` | 3 further probes | | **save incrementally** — a turn timeout destroys anything held to the end | `guest-stalls.md` | `ob_probe2.py`, four iterations later | | a roster count that is not the stage's member count is **a different stage** — discard | `mission-per-record-strength.md` | — | | **read the witness first**, before interpreting any number | `remaining-ob-hunt.md` | — | Writing each lesson down did not stop the next probe from repeating it, because each probe started from nothing. `tools/re-capture/probeharness.py` makes them structural instead. ## What it provides ```python p = Probe(baseline=116) # None = accept whatever loads if not p.ok: sys.exit(p.why) # e.g. "DISCARD: roster settled at 42" print(p.summary()) # witness count + rate, read this FIRST p.log('t\tvalue') # opens a TSV, flushed on every emit while p.tick(every=15, secs=300): p.emit(...) # written and flushed immediately print(p.status()) # '' | '*** GUEST STALLED ***' | UNVALIDATED ``` * `Probe(baseline=…)` discovers the roster, **rescans up to five times**, and refuses to start if the count never reaches the baseline. * The **witness is built in**: calibrated on construction, sampled by `tick()`, reported by `status()` and `summary()`. A probe cannot forget it, and when no witness is found it says `UNVALIDATED` rather than reporting zero stalls. * `emit()` **flushes on every line** — results cannot be lost to a timeout. * `craft()`, `strengths()`, `alive()` and `heap()` provide the roster→craft link, per-record liveness and the raw heap, so a new probe writes only its own logic. ## ✅ Verified `deploy_probe.py` reimplements the per-record deployment watch on top of it, in about 40 lines against `wave7_probe.py`'s 150. First live run: ``` roster 116, definitions 14, 32 witnesses at 10/s, stalled samples 0 t= 249s deployed=40 up=0 down=1 (cum 0/7) ``` Baseline check passed, witness found and reporting, seven losses tracked, `/tmp/deploy.tsv` written incrementally. Nothing about the result is new — that is the point; the harness reproduces a known-good measurement. ## Not done The existing probes (`wave6/7`, `ob_probe2`, `ob_by_hud`, `census`, `liveness`) are **not** ported. They work, and rewriting them risks changing results that other documents cite. New probes should use the harness; old ones should be ported only when they next need a change.