Small vision models (qwen3-vl-4b) sometimes loop the OCR array, running the response into the token ceiling (finish_reason: length) and failing to parse. Layered guardrails: - Prompts: explicit "transcribe each element once, never repeat/loop, at most N, STOP when done" in the OCR, grounding and combined prompts. - Schema: hard maxItems on ocr_results/tagging_results/content_type and maxLength on text/scene — LM Studio's grammar enforces these, so a loop is grammar-bounded rather than relying on max_tokens. - Sampling: a configurable frequency_penalty (ANALYSIS_FREQUENCY_PENALTY, default 0.3) sent with each request — the decode-time lever that actually breaks loops. - Code: sanitize() collapses runaway consecutive duplicates (same normalized text + kind) so a loop that slips through still can't flood the row. Tests: schema caps present, frequency_penalty included only when nonzero, sanitize collapses consecutive repeats; config default. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
862 B
862 B