diff --git a/docs/re/mission-freeze-resume-spin.md b/docs/re/mission-freeze-resume-spin.md index a43cf3f8..fcee0d11 100644 --- a/docs/re/mission-freeze-resume-spin.md +++ b/docs/re/mission-freeze-resume-spin.md @@ -485,6 +485,16 @@ w> F8000254 XThread::Resume: host resume was refused for thread F80001E8 <- en `Fatal` or `abort` in the whole 1.1 MB log. No shutdown line either. The process is just gone. +**And on the next occurrence the shell named it: `Killed`.** Run 8 died 54 s into +its *boot*, and `launch_mission.sh` printed + +``` +line 74: 176880 Killed nohup run-canary --apu=sdl --log_mask=13 ... +``` + +which is bash reporting **SIGKILL**. So this is not an internal fault at all — +something outside the process is killing it. + ### Memory pressure is a suspect, and only a suspect The container's cgroup, read immediately after, with **no emulator running**: @@ -505,6 +515,17 @@ disc-wide format sweeps, which read every `.pak` — was most of it. 🔴 **But *what* did. Recorded as an unexplained third failure mode rather than an OOM story, because the counter that would have proved OOM says zero. +🔴 **Checked again immediately after run 8's SIGKILL, and it is still not OOM.** +`oom_kill` remained **0** and the `max` (allocation-stall) counter did **not +move** from 4 421 — so during run 8 the cgroup never even reached its limit, +`memory.current` being 5.35 GB of 7 GiB. The host had **13.8 GB available** when +checked. Two kills, no OOM evidence either time. + +**So the cause is genuinely unidentified**, and the next occurrence is now +instrumented rather than reconstructed: `freeze_watch.sh` samples host +`MemAvailable`, the cgroup's `memory.current` and its `oom_kill` counter on every +poll, and dumps the last five samples when it sees the process disappear. + **Hygiene that follows either way:** `/dev/shm/xenia_memory_*` survives a dead run (342 MB resident here) and `run-canary` only clears it at *launch*; and `vm.drop_caches` is not writable in the container (read-only `/proc/sys`), so diff --git a/tools/re-capture/freeze_watch.sh b/tools/re-capture/freeze_watch.sh index b9202ee3..69ebd4d5 100755 --- a/tools/re-capture/freeze_watch.sh +++ b/tools/re-capture/freeze_watch.sh @@ -22,9 +22,24 @@ OUT="${2:-/sylph-home/re/logs/freeze-evidence.txt}" DEADLINE=$(( SECONDS + ${3:-1800} )) : > "$OUT" +# Sample memory every poll, so a run that DIES has a contemporaneous reading +# instead of a post-mortem guess. Two runs were SIGKILLed from outside with the +# cgroup's oom_kill counter reading 0 both times, and by the time anything was +# read the pressure (if that is what it was) had passed. +mem_line() { + local avail cur + avail=$(awk '/MemAvailable/{print int($2/1024)}' /proc/meminfo) + cur=$(( $(cat /sys/fs/cgroup/memory.current 2>/dev/null || echo 0) / 1048576 )) + local oom + oom=$(awk '/^oom_kill /{print $2}' /sys/fs/cgroup/memory.events 2>/dev/null) + echo "t=${SECONDS}s host_avail=${avail}MiB cgroup=${cur}MiB oom_kill=${oom}" +} + while [ $SECONDS -lt $DEADLINE ]; do + mem_line >> "${OUT%.txt}-mem.log" if ! pgrep -x xenia_canary >/dev/null; then - echo "EMULATOR GONE at ${SECONDS}s" | tee -a "$OUT"; exit 4 + { echo "EMULATOR GONE at ${SECONDS}s"; echo "last memory samples:"; + tail -5 "${OUT%.txt}-mem.log"; } | tee -a "$OUT"; exit 4 fi if python3 "$SD/frozen.py" 5 >/dev/null; then if python3 -c "import sys; sys.path.insert(0,'$SD'); import frozen; sys.exit(0 if frozen.in_flight() else 1)"; then