feat(analysis): gate page leasing on a vision readiness probe

The analysis worker leases an analyze_page job (incrementing attempts in
SQL) before calling vision, so a stopped/loading vision burns all retries
and writes a permanent failed page_analysis row. With the vision-manager
autoscaler idle-stopping the container, that would poison the queue on
every cold start.

Add an optional VisionReadiness seam (HTTP GET /health in production) wired
via the new env-only ANALYSIS_VISION_HEALTH_URL. When set, the worker parks
without leasing until the probe answers 2xx; jobs stay pending with
attempts untouched and resume the instant vision is ready. None preserves
today's behavior for always-on endpoints.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
MechaCat02
2026-06-14 14:57:48 +02:00
parent 2b7a11b480
commit b86aa80c87
6 changed files with 258 additions and 1 deletions

View File

@@ -78,6 +78,7 @@ fn spawn_with(
events: std::sync::Arc::new(
mangalord::analysis::events::AnalysisEvents::new(),
),
readiness: None,
},
);
(handle, cancel)
@@ -227,6 +228,7 @@ async fn worker_publishes_started_and_completed_events(pool: PgPool) {
workers: 1,
job_timeout: Duration::from_secs(5),
events: events.clone(),
readiness: None,
},
);
@@ -277,6 +279,7 @@ async fn worker_publishes_failed_event_on_dispatch_error(pool: PgPool) {
workers: 1,
job_timeout: Duration::from_secs(5),
events: events.clone(),
readiness: None,
},
);