Survive a first full crawl: poll, time out on silence, retry downloads
The first full crawl with the file manager ran 14 minutes, downloading every file once, and broke in three ways: - `schulcloud refresh` reported "fetch failed" for a crawl that was succeeding: Node's fetch abandons a response without headers after five minutes. POST /api/refresh takes wait:false and the CLI polls /api/status; refresh_index answers after 50 s and leaves the crawl running, and index_status says when a first crawl is under way. - Downloads were bounded by the 30 s request timeout, which cut 11 MB scans off mid-transfer. They now time out on 30 s of silence instead. - Failures were recorded once and never retried. A download failure is now retried on the next crawl while an extraction failure stays final, and PDF text containing NUL, which Postgres refuses, is stripped. On the re-crawl all six failed files succeeded; only two videos above the mirror cap stay metadata-only, by design. 137 tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -49,6 +49,10 @@ alongside courses, with their files under the room's name rather than a course's
|
||||
`refresh` asks the server to re-read Schulcloud. Pass `--course` when you know
|
||||
what changed: that is a handful of requests, where a full re-crawl reads every
|
||||
course. The server refuses a repeat within a minute unless you pass `--force`.
|
||||
A full re-crawl can take many minutes — the first one downloads every file,
|
||||
file-manager folders included — so `refresh` starts it and then polls the
|
||||
server's status, printing a note every half minute, rather than holding one
|
||||
request open (which Node's fetch abandons after five minutes).
|
||||
|
||||
### The file manager (`fs`)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user