Survive a first full crawl: poll, time out on silence, retry downloads

The first full crawl with the file manager ran 14 minutes, downloading every
file once, and broke in three ways:

- `schulcloud refresh` reported "fetch failed" for a crawl that was
  succeeding: Node's fetch abandons a response without headers after five
  minutes. POST /api/refresh takes wait:false and the CLI polls /api/status;
  refresh_index answers after 50 s and leaves the crawl running, and
  index_status says when a first crawl is under way.
- Downloads were bounded by the 30 s request timeout, which cut 11 MB scans
  off mid-transfer. They now time out on 30 s of silence instead.
- Failures were recorded once and never retried. A download failure is now
  retried on the next crawl while an extraction failure stays final, and PDF
  text containing NUL, which Postgres refuses, is stripped.

On the re-crawl all six failed files succeeded; only two videos above the
mirror cap stay metadata-only, by design. 137 tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
MechaCat02
2026-09-16 20:19:16 +02:00
parent bed3923902
commit 3e44e66dde
11 changed files with 310 additions and 28 deletions

View File

@@ -184,6 +184,30 @@ describe('Store', { skip: DB_URL ? false : 'set TEST_DATABASE_URL to run' }, ()
assert.ok(hits.some((h) => h.nodeId === 'f9'), 'PDF contents should be searchable, not just the filename');
});
it('stores extracted text containing NUL, which Postgres text refuses', async () => {
await store.recordFileText({
fileId: 'f9', name: 'skript.pdf', mimeType: 'application/pdf', size: 99,
content: 'Webserver\u0000 und Proxy', note: 'ok\u0000', mirrorPath: 'Info/Board/skript.pdf', mirrorSize: 99,
});
const hits = await store.search('Proxy');
assert.ok(hits.some((h) => h.nodeId === 'f9'), 'text with NUL bytes should still be stored and searchable');
});
it('keeps a file whose download failed queued for the next crawl, but not one that failed to extract', async () => {
await store.recordFileText({
fileId: 'f9', name: 'skript.pdf', mimeType: 'application/pdf', size: 99,
content: null, note: 'download failed, retried on the next crawl: timeout',
mirrorPath: null, mirrorSize: null, retry: true,
});
assert.ok((await store.filesNeedingText()).some((f) => f.fileId === 'f9'), 'a transient failure must be retried');
await store.recordFileText({
fileId: 'f9', name: 'skript.pdf', mimeType: 'application/pdf', size: 99,
content: null, note: 'extraction failed: bad xref', mirrorPath: null, mirrorSize: null,
});
assert.ok(!(await store.filesNeedingText()).some((f) => f.fileId === 'f9'), 'a parser failure is final');
});
it('resolves a timestamp cursor to a generation', async () => {
const id = await store.resolveCursor(new Date().toISOString());
assert.ok(id && id > 0);