The first full crawl with the file manager ran 14 minutes, downloading every file once, and broke in three ways: - `schulcloud refresh` reported "fetch failed" for a crawl that was succeeding: Node's fetch abandons a response without headers after five minutes. POST /api/refresh takes wait:false and the CLI polls /api/status; refresh_index answers after 50 s and leaves the crawl running, and index_status says when a first crawl is under way. - Downloads were bounded by the 30 s request timeout, which cut 11 MB scans off mid-transfer. They now time out on 30 s of silence instead. - Failures were recorded once and never retried. A download failure is now retried on the next crawl while an extraction failure stays final, and PDF text containing NUL, which Postgres refuses, is stripped. On the re-crawl all six failed files succeeded; only two videos above the mirror cap stay metadata-only, by design. 137 tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
146 lines
6.3 KiB
Markdown
146 lines
6.3 KiB
Markdown
# The `schulcloud` CLI
|
|
|
|
Browses and mirrors your Schulcloud files from a laptop, by talking to the
|
|
schulcloud-mcp server on the Pi.
|
|
|
|
## Why it goes through the Pi
|
|
|
|
The CLI never talks to Schulcloud. It holds no `jwt` cookie, no Schulcloud
|
|
credential of any kind — only this server's bearer token.
|
|
|
|
That is not an accident of layering; it solves a real problem. A Schulcloud
|
|
session dies after two hours of inactivity, and a CLI process lives for seconds,
|
|
so a CLI with its own token would be dead most times you reached for it. The Pi
|
|
already keeps one session alive around the clock. Routing through it means one
|
|
session, one keepalive, and one place to paste a fresh cookie once a month.
|
|
|
|
It also means the laptop cannot accidentally end the server's session: nothing
|
|
here can call logout.
|
|
|
|
## Setup
|
|
|
|
```bash
|
|
schulcloud login --server https://mcp.example.org --token <MCP_AUTH_TOKEN> --dir ~/Schulcloud
|
|
```
|
|
|
|
The token is the same `MCP_AUTH_TOKEN` the Claude connector uses — one token
|
|
guards both surfaces. `login` verifies it before saving, so a typo fails
|
|
immediately rather than on first real use. Config is written to
|
|
`~/.config/schulcloud/config.json` with mode `0600`.
|
|
|
|
`SCHULCLOUD_SERVER`, `SCHULCLOUD_TOKEN` and `SCHULCLOUD_SYNC_DIR` override the
|
|
file, for CI or one-off invocations.
|
|
|
|
## Commands
|
|
|
|
```
|
|
schulcloud status how fresh the server's index is
|
|
schulcloud ls [--course <id>] [--long]
|
|
schulcloud get <fileId> [--out <path>]
|
|
schulcloud sync [--dry-run] [--full] [--prune] [--dir <path>] [--jobs <n>]
|
|
schulcloud refresh [--course <id>] [--force]
|
|
```
|
|
|
|
`ls --long` prints file ids, which is what `get` takes.
|
|
|
|
`--course` accepts a **course or a room id** — rooms ("Räume") are mirrored
|
|
alongside courses, with their files under the room's name rather than a course's.
|
|
|
|
`refresh` asks the server to re-read Schulcloud. Pass `--course` when you know
|
|
what changed: that is a handful of requests, where a full re-crawl reads every
|
|
course. The server refuses a repeat within a minute unless you pass `--force`.
|
|
A full re-crawl can take many minutes — the first one downloads every file,
|
|
file-manager folders included — so `refresh` starts it and then polls the
|
|
server's status, printing a note every half minute, rather than holding one
|
|
request open (which Node's fetch abandons after five minutes).
|
|
|
|
### The file manager (`fs`)
|
|
|
|
The Schulcloud file manager ("Dateien") — Persönliche, Kurs-, Team- and
|
|
Geteilte Dateien — browsed like a filesystem, live:
|
|
|
|
```
|
|
schulcloud fs ls [path] [--long]
|
|
schulcloud fs tree [path] [--depth <n>] [--max-folders <n>]
|
|
schulcloud fs find <name> [--path <path>] [--type file|folder] [--long]
|
|
schulcloud fs get <path> [--out <path>] [--force] [--jobs <n>]
|
|
```
|
|
|
|
```console
|
|
$ schulcloud fs ls /courses
|
|
$ schulcloud fs tree "/courses/FIA24B - SK (Rh)"
|
|
$ schulcloud fs find "*Erben*" --path /courses
|
|
$ schulcloud fs get "/courses/FIA24B - SK (Rh)/02_Erbrecht" --out ~/Erbrecht
|
|
```
|
|
|
|
The tree is `/my`, `/courses/<course>`, `/teams/<team>` and `/shared`; the
|
|
German names ("/Kurs-Dateien") work too. Names may contain `/` — course names
|
|
often do — and still resolve; any segment may also be an id from `--long`.
|
|
|
|
`fs find` matches any part of a name, or, given `*` or `?`, the whole name as
|
|
`find -name` does. `fs get` on a folder downloads everything below it, keeps
|
|
the structure, and skips files already present at the same size — so re-running
|
|
it resumes. Each folder is one page load on the server, so large trees take a
|
|
while.
|
|
|
|
`sync` mirrors these files too, under `<course>/Kurs-Dateien/…`,
|
|
`Persönliche Dateien/…`, `Team-Dateien/<team>/…` and `Geteilte Dateien/`, once
|
|
the server's index includes them (`INDEX_FILE_MANAGER`, on by default).
|
|
|
|
## How sync works
|
|
|
|
It is a **one-way mirror, not a two-way sync**, and that follows from the data
|
|
rather than from laziness: Schulcloud file records are immutable — editing a
|
|
file upstream produces a *new* record — so there is no content versioning, no
|
|
conflict resolution and no merge. "Download what I do not have" is the whole
|
|
algorithm.
|
|
|
|
Local state lives in `.schulcloud-sync.json` at the root of the sync directory,
|
|
**keyed by file record id with the path as derived output**. That is what makes
|
|
renames cheap: when a teacher renames a board column, the file moves on disk
|
|
instead of being downloaded again under a new name and left duplicated under the
|
|
old one.
|
|
|
|
What it checks, and why only that:
|
|
|
|
- **Size**, not a checksum. The download endpoint exposes no `ETag` and
|
|
Schulcloud publishes no hash, so verifying content would mean re-downloading
|
|
every file to learn what it already told us. Size reliably catches the failure
|
|
that actually happens — a truncated or interrupted download — and costs a
|
|
`stat`.
|
|
- Downloads land on a `.part` neighbour and are renamed into place, so an
|
|
interrupted run never leaves a half-file that a later run mistakes for
|
|
complete.
|
|
|
|
**Deletions are not propagated by default.** A teacher removing a worksheet is
|
|
not a reason to destroy your copy of it; `sync` reports those as "gone upstream,
|
|
kept". Pass `--prune` to actually delete them.
|
|
|
|
`--dry-run` prints exactly what would happen, writes nothing, and does not
|
|
advance the cursor.
|
|
|
|
## Cursors
|
|
|
|
The server's sync cursor is a **crawl generation id**, not a timestamp. This is
|
|
deliberate and measured: `GET /course-rooms/{id}/board` returns the *request
|
|
time* as `updatedAt` for most elements, so a timestamp cursor would report every
|
|
board as changed on every crawl. Comparing generations by identity also detects
|
|
deletions, which no timestamp scheme can.
|
|
|
|
`--since` on the server API accepts an ISO date for convenience, resolved to the
|
|
nearest generation — but correctness never depends on it.
|
|
|
|
If the server no longer recognises your stored cursor it returns `409` rather
|
|
than silently treating everything as new, so you are never tricked into
|
|
re-downloading the world. Run `sync --full` deliberately in that case.
|
|
|
|
## Paths
|
|
|
|
Mirror paths are `Course/Board/Card/filename`, built by `core/paths.ts`.
|
|
|
|
Every component of that path originates in Schulcloud — course titles, card
|
|
titles and filenames are all user-supplied upstream — so each is reduced to a
|
|
single safe path component, and the result is re-checked against the sync root
|
|
before anything is written. A file named `../../.ssh/authorized_keys` cannot
|
|
escape, and `sync` refuses such an entry rather than writing it.
|