feat(analysis): position-aware OCR seam dedup for sliced pages
Boundary text duplicated across slices (often with slightly different, cropped transcriptions) survived the text-only merge. The OCR pass now asks for a per-piece vertical position and the merge uses it: - OcrResult gains optional `y` (fraction 0..1 of the slice); the Pass-A OCR schema/prompt request it (combined path leaves it None). Not persisted. - merge_ocr takes each slice's working-image y-band and maps `y` to a page-global position. A seam pair is a duplicate when position-close (same kind) OR text-similar, so a mis-OCR'd boundary line is caught even when the text differs. The kept copy is the one more central in its slice (less cropped); falls back to keep-longer when positions are missing. - Self-calibrates the model's y values (pixels vs. fraction) and ignores degenerate columns, so a bad localizer can't over-merge. Tests: position pairs differing texts and keeps the less-cropped copy; text-only fallback (dedup/keep-longer/non-adjacent/order) still holds. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -434,6 +434,7 @@ async fn seed_partial(pool: &PgPool) -> (uuid::Uuid, uuid::Uuid, uuid::Uuid, uui
|
||||
ocr_results: vec![OcrResult {
|
||||
text: "Hello".into(),
|
||||
kind: "speech".into(),
|
||||
y: None,
|
||||
}],
|
||||
tagging_results: vec!["action".into(), "city".into()],
|
||||
scene_description: "A rainy street.".into(),
|
||||
|
||||
@@ -47,6 +47,7 @@ fn analysis(
|
||||
.map(|(text, kind)| OcrResult {
|
||||
text: (*text).to_string(),
|
||||
kind: (*kind).to_string(),
|
||||
y: None,
|
||||
})
|
||||
.collect(),
|
||||
tagging_results: tags.iter().map(|t| t.to_string()).collect(),
|
||||
|
||||
@@ -54,6 +54,7 @@ fn analysis(
|
||||
.map(|(t, k)| OcrResult {
|
||||
text: (*t).into(),
|
||||
kind: (*k).into(),
|
||||
y: None,
|
||||
})
|
||||
.collect(),
|
||||
tagging_results: auto_tags.iter().map(|t| t.to_string()).collect(),
|
||||
|
||||
Reference in New Issue
Block a user