search now queries the Postgres index and states its freshness in every result, with fresh=true bypassing it for a live crawl — the agent can always get current data rather than being quietly misled by a stale index. Adds refresh_index (per-course by default; a full crawl is ~270 requests), what_changed (generation diff — the API has no changed-since filter of any kind), and index_status. /api gives the CLI its backend behind the same bearer token as /mcp: GET /manifest (cursor + per-file status), GET /files/:id (served from the mirror with Range support, falling back to a live proxy for files too large to mirror), GET /status, POST /refresh. Bytes go over plain HTTP rather than MCP because base64 in JSON-RPC costs a third more and buffers whole files. An unresolvable manifest cursor returns 409 rather than silently meaning "everything is new", so a client cannot be tricked into a full re-download. Verified end to end against the live instance and a real Postgres: crawl -> index -> German FTS -> manifest -> ranged download, with 401 on missing token, 400 on a malformed id, and 429 on a too-soon refresh. Two findings worth recording. The build silently omitted the .sql migrations from dist, which the store's graceful degradation turned into "running without the index" rather than a crash — now copied by a build step. And 3 of 4 sampled course PDFs have no embedded fonts at all: they are scans, so extraction legitimately yields nothing. That is now detected and reported as image-only with OCR named as the missing piece, instead of an indistinguishable "0 characters". It revises the roadmap's "OCR not needed" note, which held for reading images but not for indexing them. 49 tests pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
4.3 KiB
4.3 KiB