Files
Mangalord/VISION-MEMORY-YIELD.md
MechaCat02 0a9e727530 feat(vision-manager): memory-pressure yield gate
Add a fourth gate to the vision autoscaler poll loop: when host memory
gets tight, proactively SIGTERM-stop the mangalord-vision container so a
transient spike elsewhere (a crawl, CI, a backend burst) can finish
instead of the kernel OOM-killer shooting something stateful.

- read_mem_used_pct: host-wide used% from /proc/meminfo MemAvailable
  (not MemFree), ERR sentinel when MemTotal or MemAvailable is missing.
- mem_check_and_stop: active stop over MEM_HIGH_WATERMARK_PCT + arm a
  restart cooldown; only inspects docker once already over the mark.
- mem_yield_active: start inhibit while cooldown active or used% >=
  MEM_LOW_WATERMARK_PCT (two watermarks = hysteresis).
- Inhibit reuses the existing idle path via pending=0; flip-only logging;
  stop_vision gains a reason arg so idle/MAX_UPTIME/memory-yield stops are
  distinctly logged.
- Faster MEM_POLL_INTERVAL sub-poll between backlog ticks to catch spikes.

Backend verifications V1-V4 (see VISION-MEMORY-YIELD.md) all pass against
the current queue, so no backend changes are needed. RESPECT_CRAWL_MUTEX
is kept for now. New VISION_MEM_* knobs wired through docker-compose.yml
and documented in .env.example.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 14:33:34 +02:00

11 KiB
Raw Blame History

Vision-manager — memory-pressure yield (design outline)

Status: implemented in vision-manager/manager.sh (feat/vision-mem-yield). Verifications V1V4 (§5) all passed against the current backend — expired leases are reclaimed continuously on every claim poll (not boot-only), connection errors requeue with backoff, the readiness gate parks before leasing, and admin re-enqueue recovers any straggler — so no backend changes were required. RESPECT_CRAWL_MUTEX is kept for now (retiring it is deferred until the watermarks are tuned on the real box). Extends: VISION-AUTOSCALE.md (the vision-manager sidecar) and vision-manager/manager.sh. One-liner: let the manager proactively stop mangalord-vision when host memory gets dangerously tight, so a transient spike elsewhere (a Chromium crawl, a CI build, a backend burst) can complete instead of the kernel OOM-killer shooting something stateful (postgres / backend).


1. Motivation

llama-server has no idle-unload and pins ~4.4 GiB while up. On the 8 GiB Pi the existing autoscaler already stops it when the backlog is empty and (new) when analysis is disabled. Neither reacts to actual memory pressure:

  • mem_limit: 6g on the vision service caps only vision's own growth. It does nothing about aggregate host pressure across all services.
  • The static RESPECT_CRAWL_MUTEX only defers a start while a crawl runs — it can't stop a vision that's already running, and it's blind to every non-crawl memory consumer (CI build, backend spike, a second crawl worker).
  • The only backstop under true pressure today is the kernel OOM killer, which may pick postgres or the backend (long-lived, stateful, painful) rather than the one process that is cheap and safe to lose.

A manager-driven stop is the friendly version: it picks the right, restartable victim and does it gracefully (SIGTERM) before the kernel does something violent.

This is the dynamic generalization of RESPECT_CRAWL_MUTEX: "back off when RAM is actually tight" strictly dominates "never co-run with a crawl." If this ships, RESPECT_CRAWL_MUTEX can likely be retired.


2. Feasibility (confirmed)

/proc/meminfo inside the manager container reports host-wide memory (it is not namespaced), so the manager already sees the truth with zero new privilege — no socket-proxy docker info, no host mount. Verified live from inside vision-manager:

MemTotal:     8256576 kB   (7.87 GiB)
MemAvailable: 1754384 kB   → ~79% used, 21% available   (vision up + 560 queued)

Use MemAvailable, never MemFree. Linux spends "free" RAM on page cache by design, so a MemFree-based "used %" sits near 90 % during normal operation and would false-trigger constantly. MemAvailable already nets out reclaimable cache. The honest metric is:

used_pct = 100 * (1 - MemAvailable / MemTotal)

Read it in bash with awk over /proc/meminfo (both values are in kB):

read_mem_used_pct() {
  awk '/^MemTotal:/{t=$2} /^MemAvailable:/{a=$2}
       END { if (t>0) printf "%d", 100*(1-a/t); else print "ERR" }' /proc/meminfo
}

3. Behaviour

A third gate in the existing poll loop, evaluated before the backlog/enable start/stop logic:

  1. Stop gate (active). If used_pct >= HIGH_WATERMARK (e.g. 92) and vision is running → docker stop mangalord-vision, record a cooldown deadline, log loudly. This is the new capability the crawl mutex never had.
  2. Start inhibit. While used_pct >= LOW_WATERMARK (e.g. 80) or a cooldown is in effect → refuse to start vision even if backlog > 0 (same shape as the crawl-mutex deferral: log once per episode, not every tick).
  3. Resume. Once used_pct < LOW_WATERMARK and the cooldown has elapsed, normal autoscaler behaviour resumes (start on backlog, idle-stop, etc.).

Hysteresis is mandatory (two distinct watermarks). A single threshold oscillates around the boundary.

The cooldown is mandatory and is the subtle part. If vision itself is the 4.4 GiB hog, stopping it drops pressure → manager sees backlog → restarts it → pressure climbs → stop … infinite thrash, each cycle paying the ~minutes cold reload. The cooldown (refuse-to-restart for N minutes after a pressure-stop, regardless of backlog) is what gives the other workload — the one that actually needed the RAM — room to finish before vision muscles back in. The yield only makes sense when something else needs the memory; the cooldown encodes that.

Gate ordering in the loop

loop every POLL_INTERVAL:
    used = read_mem_used_pct()
    pending = pending_analysis()
    if not analysis_enabled():        pending = 0          # existing gate
    if mem_yield_active(used):         pending = 0          # NEW: inhibit starts
    if used >= HIGH_WATERMARK and vision running:           # NEW: active stop
        stop_vision(); start_cooldown(); continue
    ... existing start-on-backlog / idle-debounce-stop logic on `pending` ...

Note the symmetry with the analysis-disabled gate already in manager.sh: forcing pending = 0 reuses the existing idle/stop path for the inhibit case, while the explicit stop_vision() handles the active case the idle path can't.


4. Proposed configuration (all env, matching manager.sh style)

Var Default Meaning
MEM_YIELD_ENABLED 1 Master switch for the whole feature.
MEM_HIGH_WATERMARK_PCT 92 used_pct ≥ this → active-stop a running vision.
MEM_LOW_WATERMARK_PCT 80 used_pct ≥ this → inhibit starts (no restart until below).
MEM_YIELD_COOLDOWN 300 Seconds after a pressure-stop during which restart is refused regardless of backlog.
MEM_POLL_INTERVAL POLL_INTERVAL Optional faster cadence for the mem check (pressure can spike between 20 s ticks).

Watermarks want tuning on the real box; 92/80 are starting points. Keep HIGH LOW ≥ ~10 so the band is wider than normal jitter.


5. Required verifications (do these BEFORE building)

The queue is built for expendability — crawler_jobs is a lease queue:

state | attempts | max_attempts(=5) | leased_until | last_error | scheduled_at

A graceful docker stop (SIGTERM → clean llama-server exit) means at most the one in-flight analyze_page job is affected; the readiness gate (ANALYSIS_VISION_HEALTH_URL) then parks the worker so it won't hammer the dead endpoint or burn the other queued jobs. But two things must be confirmed in backend/src/analysis/daemon.rs (not visible from the DB schema alone):

  • V1 — reclaim cadence. The bugfix/crawler-recovery-hardening branch added reclaim-orphaned at startup. Confirm expired leases are also reclaimed on the normal claim poll (i.e. the claim query includes something like OR (state='running' AND leased_until < now())). If reclaim is boot-only, a pressure-kill strands that one job as running with a dead lease until the next backend restart → the page silently never analyses. This decides how aggressive we can safely be.
  • V2 — error path requeue. When the worker's HTTP call to vision is refused (vision killed), confirm the job is returned to pending (attempts++) and not marked permanently failed on a connection error. A connection-refused/transport error should be retryable, distinct from a model-returned bad response.
  • V3 — park does not deadlock. Confirm that when the readiness gate parks the worker, it is not holding an un-renewed lease that blocks the job indefinitely (it should either release on park or let the lease expire + be reclaimed per V1).
  • V4 — attempts accounting. A pressure-kill is "our fault," yet still burns one of the 5 attempts. Confirm repeated yields on the same in-flight page can't quietly exhaust retries → permanent failed. Mitigations if it's a concern: bump max_attempts, or rely on the admin re-enqueue (all/manga/ chapter) you already have to recover stragglers.

Empirical check (once V1V3 look right): with analysis on and one job running, docker stop mangalord-vision, then watch SELECT state, attempts, last_error FROM crawler_jobs WHERE state IN ('running','pending') AND payload->>'kind'='analyze_page' — the killed job should reappear as pending with attempts incremented, and resume when vision restarts. This is the acceptance test for the whole feature.


6. Pitfalls

  • MemFree vs MemAvailable — using MemFree makes it fire constantly (page cache). Already addressed by §2, but it's the #1 way to get this wrong.
  • Thrash / feedback loop — if vision is the hog, stop→restart oscillates. The cooldown (§3) is the guard; without it the feature is worse than nothing.
  • Spikes between ticks — a Chromium burst can blow past 92 % inside a 20 s poll gap and the kernel OOM-kills before the manager's next read. The valve reduces OOM risk, it doesn't eliminate it. Consider a shorter MEM_POLL_INTERVAL for just the mem read, and treat this as defence-in-depth alongside (not instead of) mem_limit and swap/zram headroom.
  • Measuring the wrong scope — confirm /proc/meminfo in the container keeps showing host totals under this kernel/cgroup setup (it does today: MemTotal 8256576 kB == the Pi). If a future runtime namespaces it, the metric silently becomes the container's tiny cgroup and the gate never/always fires.
  • Cold-start cost — every pressure-stop pays ~minutes to mmap + warm 3.2 GiB on the next start. Don't set the watermarks so tight that normal operation trips them; the backlog will stall during each reload.
  • Interaction with the idle-stop and disable gate — make sure the three reasons vision can be down (idle-drained, analysis-disabled, memory-yield) are distinctly logged (flip-only, like the existing gates) so a stop is never misdiagnosed. A single muddy "stopped vision" line will cost debugging time.
  • Double-stop / start racesdocker stop then an immediate backlog-driven docker start in the same or next tick. The cooldown plus checking vision_running before acting prevents this; don't drop those guards.
  • OOM-killer still owns the tailmem_limit: 6g caps vision's own growth but the host can still be pushed over by the sum of everything else. This feature lowers the probability of a bad OOM kill; pairing it with a little swap or zram gives the kernel a cushion to survive the gap between a spike and the manager's reaction.

7. Recommendation

Worth building, as a third gate in the existing manager.sh poll loop: MemAvailable-based metric, two watermarks (hysteresis), a restart cooldown, an active stop (the capability the crawl mutex lacks), and distinct flip-only logging. It supersedes RESPECT_CRAWL_MUTEX (which can then be removed).

Gate it on verification V1 first — whether expired leases are reclaimed continuously or only at boot decides whether a yielded job cleanly retries or strands. If V1/V2 hold, this is a low-risk, high-value safety valve; if reclaim is boot-only, fix that in the backend before enabling memory-yield, otherwise each pressure-stop quietly drops a page until the next backend restart.

Keep the manager dumb: this is ~30 lines of bash (one awk reader, two watermark comparisons, one cooldown timestamp) — no new dependencies, no new privilege, no web surface.