feat(analysis): position-aware OCR seam dedup for sliced pages

Boundary text duplicated across slices (often with slightly different,
cropped transcriptions) survived the text-only merge. The OCR pass now
asks for a per-piece vertical position and the merge uses it:

- OcrResult gains optional `y` (fraction 0..1 of the slice); the Pass-A
  OCR schema/prompt request it (combined path leaves it None). Not
  persisted.
- merge_ocr takes each slice's working-image y-band and maps `y` to a
  page-global position. A seam pair is a duplicate when position-close
  (same kind) OR text-similar, so a mis-OCR'd boundary line is caught even
  when the text differs. The kept copy is the one more central in its slice
  (less cropped); falls back to keep-longer when positions are missing.
- Self-calibrates the model's y values (pixels vs. fraction) and ignores
  degenerate columns, so a bad localizer can't over-merge.

Tests: position pairs differing texts and keeps the less-cropped copy;
text-only fallback (dedup/keep-longer/non-adjacent/order) still holds.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
MechaCat02
2026-06-13 23:17:55 +02:00
parent 85d65f5eda
commit 25aba3ac58
9 changed files with 153 additions and 44 deletions

View File

@@ -434,6 +434,7 @@ async fn seed_partial(pool: &PgPool) -> (uuid::Uuid, uuid::Uuid, uuid::Uuid, uui
ocr_results: vec![OcrResult {
text: "Hello".into(),
kind: "speech".into(),
y: None,
}],
tagging_results: vec!["action".into(), "city".into()],
scene_description: "A rainy street.".into(),

View File

@@ -47,6 +47,7 @@ fn analysis(
.map(|(text, kind)| OcrResult {
text: (*text).to_string(),
kind: (*kind).to_string(),
y: None,
})
.collect(),
tagging_results: tags.iter().map(|t| t.to_string()).collect(),

View File

@@ -54,6 +54,7 @@ fn analysis(
.map(|(t, k)| OcrResult {
text: (*t).into(),
kind: (*k).into(),
y: None,
})
.collect(),
tagging_results: auto_tags.iter().map(|t| t.to_string()).collect(),