Index-backed search, /api surfaces, image-only PDF detection
search now queries the Postgres index and states its freshness in every result, with fresh=true bypassing it for a live crawl — the agent can always get current data rather than being quietly misled by a stale index. Adds refresh_index (per-course by default; a full crawl is ~270 requests), what_changed (generation diff — the API has no changed-since filter of any kind), and index_status. /api gives the CLI its backend behind the same bearer token as /mcp: GET /manifest (cursor + per-file status), GET /files/:id (served from the mirror with Range support, falling back to a live proxy for files too large to mirror), GET /status, POST /refresh. Bytes go over plain HTTP rather than MCP because base64 in JSON-RPC costs a third more and buffers whole files. An unresolvable manifest cursor returns 409 rather than silently meaning "everything is new", so a client cannot be tricked into a full re-download. Verified end to end against the live instance and a real Postgres: crawl -> index -> German FTS -> manifest -> ranged download, with 401 on missing token, 400 on a malformed id, and 429 on a too-soon refresh. Two findings worth recording. The build silently omitted the .sql migrations from dist, which the store's graceful degradation turned into "running without the index" rather than a crash — now copied by a build step. And 3 of 4 sampled course PDFs have no embedded fonts at all: they are scans, so extraction legitimately yields nothing. That is now detected and reported as image-only with OCR named as the missing piece, instead of an indistinguishable "0 characters". It revises the roadmap's "OCR not needed" note, which held for reading images but not for indexing them. 49 tests pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -53,3 +53,51 @@ describe('formatBytes', () => {
|
||||
assert.equal(formatBytes(5 * 1024 * 1024), '5.0 MB');
|
||||
});
|
||||
});
|
||||
|
||||
/**
|
||||
* Builds a structurally valid PDF with correct xref offsets — pdfjs rejects
|
||||
* anything less, so a hand-waved byte string would test the error path instead
|
||||
* of the one we care about.
|
||||
*/
|
||||
function minimalPdf(options: { withFont: boolean }): Buffer {
|
||||
const content = options.withFont
|
||||
? 'BT /F1 12 Tf 10 100 Td (Hallo Welt) Tj ET'
|
||||
: 'q 100 0 0 100 10 10 cm /Im0 Do Q'; // draws an image, no text operators
|
||||
const objs = [
|
||||
'<</Type/Catalog/Pages 2 0 R>>',
|
||||
'<</Type/Pages/Kids[3 0 R]/Count 1>>',
|
||||
`<</Type/Page/Parent 2 0 R/MediaBox[0 0 200 200]/Resources<<${
|
||||
options.withFont ? '/Font<</F1 5 0 R>>' : '/XObject<</Im0 5 0 R>>'
|
||||
}>>/Contents 4 0 R>>`,
|
||||
`<</Length ${content.length}>>stream\n${content}\nendstream`,
|
||||
options.withFont ? '<</Type/Font/Subtype/Type1/BaseFont/Helvetica>>' : '<</Subtype/Image/Width 1/Height 1>>',
|
||||
];
|
||||
let out = '%PDF-1.4\n';
|
||||
const offsets: number[] = [];
|
||||
objs.forEach((body, index) => {
|
||||
offsets.push(out.length);
|
||||
out += `${index + 1} 0 obj${body}endobj\n`;
|
||||
});
|
||||
const xref = out.length;
|
||||
out += `xref\n0 ${objs.length + 1}\n0000000000 65535 f \n`;
|
||||
for (const offset of offsets) out += `${String(offset).padStart(10, '0')} 00000 n \n`;
|
||||
out += `trailer<</Size ${objs.length + 1}/Root 1 0 R>>\nstartxref\n${xref}\n%%EOF`;
|
||||
return Buffer.from(out, 'latin1');
|
||||
}
|
||||
|
||||
describe('image-only PDFs', () => {
|
||||
it('extracts normally when the PDF has a text layer', async () => {
|
||||
const result = await extractContent(minimalPdf({ withFont: true }), 'application/pdf', 'doc.pdf', 10_000);
|
||||
assert.equal(result.kind, 'text');
|
||||
assert.match(result.text ?? '', /Hallo Welt/);
|
||||
});
|
||||
|
||||
it('reports a missing text layer rather than a bare zero-character result', async () => {
|
||||
// Measured on the real account: 3 of 4 sampled course PDFs are image-only,
|
||||
// so "0 characters" must not look like a parser failure.
|
||||
const result = await extractContent(minimalPdf({ withFont: false }), 'application/pdf', 'scan.pdf', 10_000);
|
||||
assert.equal(result.kind, 'binary');
|
||||
assert.match(result.note, /image-only PDF/);
|
||||
assert.match(result.note, /OCR/);
|
||||
});
|
||||
});
|
||||
|
||||
Reference in New Issue
Block a user