Files
Mangalord/VISION-AUTOSCALE.md
MechaCat02 f441425519
Some checks failed
deploy / test-backend (push) Failing after 7m32s
deploy / build-and-push (push) Has been cancelled
deploy / deploy (push) Has been cancelled
deploy / test-frontend (push) Has been cancelled
fix(analysis): harden readiness gate and vision-manager per self-review
Backend:
- Log on readiness state transitions (ready<->not-ready) so a misconfigured
  ANALYSIS_VISION_HEALTH_URL, which otherwise parks the worker silently
  forever, is diagnosable.
- Add unit tests for HttpVisionReadiness: 2xx->ready, non-2xx->not-ready,
  connection error->not-ready (the production status mapping had no test).

vision-manager:
- Guard the backlog count against non-numeric output before `-gt`, so a stray
  value can't exit-2 and kill the loop under `set -e`.
- Throttle the crawl-mutex "deferring start" log to once per episode.
- Only reset the idle/uptime timers when `docker stop` actually succeeds, so a
  failed stop retries next tick instead of waiting a full debounce window.
- Decouple the per-probe curl timeout (HEALTH_TIMEOUT) from the warm-up poll
  cadence.

Docs:
- Correct the docker-socket-proxy comment: CONTAINERS+POST permits the full
  container lifecycle (not just start/stop); state the real trust boundary.
- Document that the externally-defined mangalord-vision container must share
  the compose network for name resolution.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 15:15:40 +02:00

139 lines
7.3 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Vision auto start/stop — design brief for the dev agent
**Goal:** the `mangalord-vision` (llama.cpp) container should only run while there
is analysis work, and stop (freeing its ~4 GiB) once the queue drains —
*without* handing the internet-facing backend control of the host Docker daemon.
`llama-server` has **no idle-unload** (unlike Ollama). The only lever is the
container lifecycle: start it when work appears, stop it when work is gone.
> **Status: implemented.** The sidecar lives in [vision-manager/](vision-manager/)
> (`manager.sh` + `Dockerfile` + `readonly-role.sql`), wired into
> [docker-compose.yml](docker-compose.yml) behind `profiles: [ai]` alongside a
> scoped `docker-socket-proxy`. The companion **readiness gate** is in the
> analysis worker: when `ANALYSIS_VISION_HEALTH_URL` is set the worker will not
> lease a page until vision answers `GET /health` 2xx, so an idle-stopped or
> still-loading vision never burns a job's retries or writes a `failed` row
> (the gotcha called out below). Configure via the `ai`-profile vars in
> [.env.example](.env.example).
>
> **Operator prerequisite:** the `mangalord-vision` container is defined
> *outside* this compose project (the manager only drives it by name). For the
> manager and backend to resolve it by name for the `/health` probe, attach it
> to this project's default network (`<project>_default`, e.g.
> `mangalord_default`) — or point `VISION_HEALTH_URL` /
> `ANALYSIS_VISION_HEALTH_URL` at an address that resolves. If the container is
> unreachable the manager silently treats it as "not running / not ready" and
> will loop trying to start a container it cannot see.
## Chosen approach — Option 2: a "vision-manager" sidecar
A tiny, single-purpose container that:
1. **Watches the analysis backlog** in Postgres (read-only DB user).
2. **Starts** `mangalord-vision` when there is pending work.
3. **Stops** it after the backlog has been empty for a debounce window.
The backend (`mangalord-backend`) is **not modified and gets no Docker access**
all privilege is isolated in the manager. This is the whole point: a popped
backend can't reach the Docker socket.
```
┌────────────────┐ reads pending count ┌────────────┐
│ vision-manager │ ───────────────────────▶│ postgres │
│ (has socket) │ └────────────┘
└──────┬─────────┘
start/stop one │ (raw socket, OR scoped via docker-socket-proxy)
container only ▼
┌────────────────┐
│ mangalord-vision│ (profiles: [ai], started by name)
└────────────────┘
```
### What to poll (concrete)
The analysis daemon consumes the **shared `crawler_jobs` queue, filtered by the
analysis job kind**, and every unanalyzed page has a `page_analysis` row with
`status='pending'`. Either is a valid "is there work?" signal:
- **Queue-accurate (what the manager uses):** the queue column is `state`
(not `status`) and the kind lives in the JSONB payload — see
[backend/migrations/0012_crawler.sql](backend/migrations/0012_crawler.sql):
```sql
SELECT count(*) FROM crawler_jobs
WHERE payload->>'kind' = 'analyze_page'
AND state IN ('pending','running');
```
- **Backlog-simple:** `SELECT count(*) FROM page_analysis WHERE status='pending'`.
Prefer a **read-only** DB role scoped to those tables. Don't reuse the backend's
DB credentials.
### Lifecycle logic (sketch)
```
loop every POLL_INTERVAL (e.g. 20s):
pending = count_pending_work()
if pending > 0 and vision is stopped:
docker start mangalord-vision
wait for GET /health == 200 (up to a few minutes — cold model load)
reset idle timer
if pending == 0 and vision is running:
if idle for >= STOP_DEBOUNCE (e.g. 510 min):
docker stop mangalord-vision
```
## Recommendations
- **Scope the Docker access.** Best: run `tecnativa/docker-socket-proxy` on an
internal-only network with everything disabled except container start/stop,
and point the manager at the proxy. Acceptable for a tiny trusted manager:
mount the raw socket *into the manager only* (never the backend).
- **Keep the manager dumb and small.** ~100 lines. A shell loop with
`docker`/`curl`, or a small Go/Rust binary. No web surface.
- **Make it the single source of truth** for the vision lifecycle. Don't also
have the backend or a cron poke the same container.
- **Start by container name**, not `compose up` (the service is `profiles:[ai]`;
name-based `docker start/stop` works and won't fight a separate `compose up`).
- Optionally expose the toggle as a runtime setting ("auto-manage vision: on/off")
so it can be disabled without redeploying.
## Pitfalls / gotchas (flag all of these)
- **Cold start is expensive (~minutes)** to mmap 3.2 GiB + warm up. So:
- **Debounce the STOP** (510 min idle) — bursty uploads/re-analysis enqueue
in waves; stopping the instant the queue hits zero thrashes the load cycle
and loses more time than it saves.
- **Gate the first request on `GET /health == 200`** after start, or the first
jobs fail with connection-refused. (The backend already has request/job
timeouts — 300s/1800s — but a multi-minute cold start can still exceed a
single request timeout if a job is dispatched too eagerly.)
- **Idempotency / single-flight:** starting an already-running container must be
a no-op; never issue concurrent starts. One manager instance only.
- **Leak safety / don't depend on a "drained" signal from the backend:** the
manager's own idle timer must stop the container even if the backend crashes
mid-batch. Conversely, if the manager dies, the container keeps running
(harmless, just RAM) — consider a max-uptime backstop.
- **RAM headroom on the 8 GiB Pi:** vision sits ~4.4 GiB; `mem_limit: 6g` is set
on the service. The manager must not start vision while another big consumer
(a crawl with Chromium) is running, or the box OOMs. Consider a simple mutual
exclusion / total-RAM check.
- **Health vs readiness:** `/health` returns 200 once the model is loaded; that
is the readiness signal. Don't treat "container running" as "ready".
- **Re-analysis storms:** an admin "re-enqueue all" can flood the queue; the
manager will (correctly) start vision and keep it up for a long backfill —
expected, but make sure the STOP debounce doesn't bounce it mid-backfill if
the queue briefly empties between batches.
- **Profile interaction:** `docker compose --profile ai up` / CI deploys name
only the two mangalord services, so they won't start/stop vision — but a
human running a full `compose --profile ai up -d` could. Document that the
manager owns the lifecycle.
## Why not give the backend the socket directly
`mangalord-backend` is internet-facing (behind Caddy at manga.mc02.dev). Mounting
the raw Docker socket there makes any backend RCE a full host takeover
(postgres, gitea, vaultwarden, …). If the backend *must* drive it, use the
docker-socket-proxy scoped to start/stop of the single container — never the raw
socket. Option 2 sidesteps the question entirely by keeping the backend clean.