Very tall webtoon pages were squashed by the fixed longest-edge downscale, so OCR was garbage. The vision client now slices to the model's pixel budget at native resolution and assembles one result: - plan_slices: native-resolution bands sized to max_pixels (portrait OR landscape), only when height/width > tall_aspect_threshold; min_slice_ height guards pathologically wide pages (reduce width); max_slices caps the count (coarser fallback). Normal pages still take one combined call. - Two-pass, OCR-first: Pass A OCRs each band (near-native), merge_ocr stitches them with seam-scoped fuzzy dedup (keep the longer transcript, preserve order/kind, don't collapse non-adjacent repeats); Pass B feeds the whole downscaled image + the merged OCR text to get tags/scene/safety grounded in the real dialogue. Only OCR is merged. - Config: replace max_image_dim with max_pixels / min_slice_height / slice_overlap / tall_aspect_threshold / max_slices (+ env_f64); bump job_timeout default 180->600 (N sequential slice calls per page); raise MAX_OCR_PIECES 60->200. New prompts/schemas: OCR_PROMPT/ocr_json_schema, GROUNDING_PROMPT/grounding_json_schema. Tests: plan_slices (single/portrait/landscape/pathological-wide/cap + coverage/overlap), merge_ocr (seam dedup, keep-longer, non-adjacent repeats, order), render_whole/render_slice, the new request builders, and updated config defaults/env. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2.3 KiB
2.3 KiB