Add the schulcloud CLI, and document the split
The CLI talks only to the Pi's /api surface and holds no Schulcloud credential — only the same bearer token the Claude connector uses. That is not layering for its own sake: a Schulcloud session dies after two hours idle and a CLI process lives for seconds, so a CLI with its own token would be dead most times you reached for it. Routing through the Pi means one session, one keepalive, one monthly cookie paste. sync is a one-way mirror, which follows from the data rather than from scope-cutting: file records are immutable upstream, so there is no versioning, no conflict resolution and no merge. State is keyed by file record id with the path as derived output, so an upstream rename moves the local file instead of duplicating it — verified against the live server. Verification is size-only because the download endpoint exposes no ETag and Schulcloud publishes no hash; size still catches the failure that happens, a truncated download. Downloads land on a .part neighbour and are renamed, so an interrupted run leaves no half-file that a later run mistakes for complete. Deletions are reported but not propagated — a teacher removing a worksheet is no reason to destroy the student's copy — with --prune to opt in. what_changed now clamps to the oldest stored generation instead of refusing, and says it did: "what's new this week" is a reasonable question to ask a two-day-old index. Two build bugs caught by the checks rather than by luck: the smoke harness constructed the app without services, so the index-backed tools were never exercised; and the Docker build could not see scripts/copy-assets.mjs, so the image would have shipped without migrations and silently degraded to live-only. 67 unit tests (9 needing Postgres), smoke green both ways — 34 checks with an index, 32 without, because graceful degradation is a supported mode and not a fallback nobody runs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
105
docs/CLI.md
Normal file
105
docs/CLI.md
Normal file
@@ -0,0 +1,105 @@
|
||||
# The `schulcloud` CLI
|
||||
|
||||
Browses and mirrors your Schulcloud files from a laptop, by talking to the
|
||||
schulcloud-mcp server on the Pi.
|
||||
|
||||
## Why it goes through the Pi
|
||||
|
||||
The CLI never talks to Schulcloud. It holds no `jwt` cookie, no Schulcloud
|
||||
credential of any kind — only this server's bearer token.
|
||||
|
||||
That is not an accident of layering; it solves a real problem. A Schulcloud
|
||||
session dies after two hours of inactivity, and a CLI process lives for seconds,
|
||||
so a CLI with its own token would be dead most times you reached for it. The Pi
|
||||
already keeps one session alive around the clock. Routing through it means one
|
||||
session, one keepalive, and one place to paste a fresh cookie once a month.
|
||||
|
||||
It also means the laptop cannot accidentally end the server's session: nothing
|
||||
here can call logout.
|
||||
|
||||
## Setup
|
||||
|
||||
```bash
|
||||
schulcloud login --server https://mcp.example.org --token <MCP_AUTH_TOKEN> --dir ~/Schulcloud
|
||||
```
|
||||
|
||||
The token is the same `MCP_AUTH_TOKEN` the Claude connector uses — one token
|
||||
guards both surfaces. `login` verifies it before saving, so a typo fails
|
||||
immediately rather than on first real use. Config is written to
|
||||
`~/.config/schulcloud/config.json` with mode `0600`.
|
||||
|
||||
`SCHULCLOUD_SERVER`, `SCHULCLOUD_TOKEN` and `SCHULCLOUD_SYNC_DIR` override the
|
||||
file, for CI or one-off invocations.
|
||||
|
||||
## Commands
|
||||
|
||||
```
|
||||
schulcloud status how fresh the server's index is
|
||||
schulcloud ls [--course <id>] [--long]
|
||||
schulcloud get <fileId> [--out <path>]
|
||||
schulcloud sync [--dry-run] [--full] [--prune] [--dir <path>] [--jobs <n>]
|
||||
schulcloud refresh [--course <id>] [--force]
|
||||
```
|
||||
|
||||
`ls --long` prints file ids, which is what `get` takes.
|
||||
|
||||
`refresh` asks the server to re-read Schulcloud. Pass `--course` when you know
|
||||
what changed: that is a handful of requests, where a full re-crawl reads every
|
||||
course. The server refuses a repeat within a minute unless you pass `--force`.
|
||||
|
||||
## How sync works
|
||||
|
||||
It is a **one-way mirror, not a two-way sync**, and that follows from the data
|
||||
rather than from laziness: Schulcloud file records are immutable — editing a
|
||||
file upstream produces a *new* record — so there is no content versioning, no
|
||||
conflict resolution and no merge. "Download what I do not have" is the whole
|
||||
algorithm.
|
||||
|
||||
Local state lives in `.schulcloud-sync.json` at the root of the sync directory,
|
||||
**keyed by file record id with the path as derived output**. That is what makes
|
||||
renames cheap: when a teacher renames a board column, the file moves on disk
|
||||
instead of being downloaded again under a new name and left duplicated under the
|
||||
old one.
|
||||
|
||||
What it checks, and why only that:
|
||||
|
||||
- **Size**, not a checksum. The download endpoint exposes no `ETag` and
|
||||
Schulcloud publishes no hash, so verifying content would mean re-downloading
|
||||
every file to learn what it already told us. Size reliably catches the failure
|
||||
that actually happens — a truncated or interrupted download — and costs a
|
||||
`stat`.
|
||||
- Downloads land on a `.part` neighbour and are renamed into place, so an
|
||||
interrupted run never leaves a half-file that a later run mistakes for
|
||||
complete.
|
||||
|
||||
**Deletions are not propagated by default.** A teacher removing a worksheet is
|
||||
not a reason to destroy your copy of it; `sync` reports those as "gone upstream,
|
||||
kept". Pass `--prune` to actually delete them.
|
||||
|
||||
`--dry-run` prints exactly what would happen, writes nothing, and does not
|
||||
advance the cursor.
|
||||
|
||||
## Cursors
|
||||
|
||||
The server's sync cursor is a **crawl generation id**, not a timestamp. This is
|
||||
deliberate and measured: `GET /course-rooms/{id}/board` returns the *request
|
||||
time* as `updatedAt` for most elements, so a timestamp cursor would report every
|
||||
board as changed on every crawl. Comparing generations by identity also detects
|
||||
deletions, which no timestamp scheme can.
|
||||
|
||||
`--since` on the server API accepts an ISO date for convenience, resolved to the
|
||||
nearest generation — but correctness never depends on it.
|
||||
|
||||
If the server no longer recognises your stored cursor it returns `409` rather
|
||||
than silently treating everything as new, so you are never tricked into
|
||||
re-downloading the world. Run `sync --full` deliberately in that case.
|
||||
|
||||
## Paths
|
||||
|
||||
Mirror paths are `Course/Board/Card/filename`, built by `core/paths.ts`.
|
||||
|
||||
Every component of that path originates in Schulcloud — course titles, card
|
||||
titles and filenames are all user-supplied upstream — so each is reduced to a
|
||||
single safe path component, and the result is re-checked against the sync root
|
||||
before anything is written. A file named `../../.ssh/authorized_keys` cannot
|
||||
escape, and `sync` refuses such an entry rather than writing it.
|
||||
@@ -3,12 +3,19 @@
|
||||
## The shape of it
|
||||
|
||||
```
|
||||
claude.ai ──HTTPS──▶ VPS (public IP) ──tunnel──▶ Pi 5 (home network)
|
||||
└─ Caddy ──▶ schulcloud-mcp:8080
|
||||
│
|
||||
└──▶ schulcloud-thueringen.de
|
||||
claude.ai ──HTTPS──┐
|
||||
├─▶ VPS (public IP) ──tunnel──▶ Pi 5 (home network)
|
||||
schulcloud CLI ─────┘ └─ Caddy ─▶ schulcloud-mcp:8080
|
||||
├─ /mcp (Claude)
|
||||
├─ /api (CLI)
|
||||
├─ Postgres (index)
|
||||
├─ mirror (file bytes)
|
||||
└──▶ schulcloud-thueringen.de
|
||||
```
|
||||
|
||||
Both front ends use the same hostname and the same bearer token. `/mcp` speaks
|
||||
MCP; `/api` serves the CLI's manifest, file bytes and re-crawl requests.
|
||||
|
||||
Claude's custom connectors call the endpoint from Anthropic's cloud, so it must
|
||||
be publicly reachable over real TLS — a localhost tunnel or self-signed cert
|
||||
will not do. The VPS provides the public address; Caddy on the Pi terminates
|
||||
@@ -64,6 +71,26 @@ networks:
|
||||
name: <the network name you just found>
|
||||
```
|
||||
|
||||
### Postgres
|
||||
|
||||
The index needs a database. On the Pi, use the existing PostgreSQL rather than
|
||||
the container in `docker-compose.yml` — create a database and user for it:
|
||||
|
||||
```sql
|
||||
CREATE USER schulcloud WITH PASSWORD '…';
|
||||
CREATE DATABASE schulcloud OWNER schulcloud;
|
||||
```
|
||||
|
||||
Then set `DATABASE_URL` in `.env` and delete the `postgres` service from the
|
||||
compose file. Migrations run automatically at startup; `pg_trgm` is created by
|
||||
the first migration, which needs the database owner to be able to
|
||||
`CREATE EXTENSION`.
|
||||
|
||||
Without `DATABASE_URL` the server still runs: search crawls live on every call
|
||||
and `/api` returns `503`. The startup log says which mode it is in.
|
||||
|
||||
### Caddy
|
||||
|
||||
Append `deploy/Caddyfile.snippet` to the Pi's Caddyfile, replacing
|
||||
`mcp.example.org` with the real hostname, and reload:
|
||||
|
||||
@@ -151,10 +178,16 @@ npm run probe # re-verify the API assumptions
|
||||
- **Sessions** are in-memory and dropped after 30 minutes idle. A restart
|
||||
invalidates them; Claude re-initializes transparently.
|
||||
- **Logs** are capped at 3 × 10 MB. The Authorization header is never logged.
|
||||
- **The container is read-only** with `cap_drop: ALL` and
|
||||
`no-new-privileges`, running as the unprivileged `node` user. It writes
|
||||
nothing to disk — downloads are streamed through memory, capped at
|
||||
`MAX_DOWNLOAD_BYTES` (25 MiB default).
|
||||
- **The container is read-only** with `cap_drop: ALL` and `no-new-privileges`,
|
||||
running as the unprivileged `node` user. The one writable path is the mirror
|
||||
volume at `/data/mirror`, which holds downloaded file bytes; everything else
|
||||
stays read-only.
|
||||
- **The mirror grows.** It holds a copy of every course file under
|
||||
`MIRROR_MAX_BYTES` (64 MiB default). Larger files — videos, mostly — are
|
||||
indexed as metadata and proxied live on request instead. Budget a few GB.
|
||||
- **A re-crawl of unchanged content downloads nothing**, because Schulcloud file
|
||||
records are immutable, so the 6-hourly crawl costs a few hundred cheap GETs in
|
||||
the steady state.
|
||||
- **The Schulcloud session has a 2-hour sliding TTL**, so the server calls
|
||||
`refresh-session` every 30 minutes. Watch for
|
||||
`keepalive: session extended, 7200s` in the logs, or run
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
# Possible extensions
|
||||
# Extensions: built and possible
|
||||
|
||||
**Nothing here is built.** This records options considered during setup, the
|
||||
decisions already taken, and the trade-offs found while building — so none of it
|
||||
has to be rediscovered.
|
||||
Items 1–3 are **now implemented** — see `docs/CLI.md` and the `store/`,
|
||||
`indexer/` and `cli/` modules. What remains below is the rationale (worth
|
||||
keeping) and the items still open.
|
||||
|
||||
## Decisions already taken
|
||||
|
||||
@@ -17,7 +17,7 @@ Two of the rejected options need no revisiting. OAuth 2.1 is disproportionate
|
||||
machinery for a single-user endpoint a bearer token already protects. "Raw bytes
|
||||
as well as extracted text" got built anyway — `download_file` takes `raw: true`.
|
||||
|
||||
## 1. Full-text search over file contents
|
||||
## 1. Full-text search over file contents — BUILT
|
||||
|
||||
**The gap:** `search` covers course/board/card/lesson/task titles and text, and
|
||||
file *names* — never file *contents*. The 93 PDFs in this account are opaque to
|
||||
@@ -42,7 +42,7 @@ several seconds, every time.
|
||||
finds *Verschlüsselung* without sharing a word. Needs an embedding model in the
|
||||
loop, so it is a separate step, not part of this.
|
||||
|
||||
## 2. Cache, with bypass
|
||||
## 2. Cache, with bypass — BUILT (`search fresh=true`, `refresh_index`)
|
||||
|
||||
**The split that makes staleness tolerable** — index is *discovery*, live API is
|
||||
*detail*. Search the index to find where something is; always re-fetch it to read
|
||||
@@ -63,14 +63,14 @@ list as current.
|
||||
outside Schulcloud. It does not weaken the read-only property, but it is new —
|
||||
consider disk encryption and whether it lands in a backup.
|
||||
|
||||
## 3. "What's new since …"
|
||||
## 3. "What's new since …" — BUILT (`what_changed`)
|
||||
|
||||
Falls out of having an index with history: diff successive crawls to surface new
|
||||
boards, cards, files and tasks. **Impossible today at any speed** — the API has no
|
||||
changed-since filter anywhere. Arguably the most useful item on this list for a
|
||||
student, and nearly free once the sync job exists.
|
||||
|
||||
## 4. Smaller items
|
||||
## 4. Still open
|
||||
|
||||
- **Video/audio transcription** — this account has 5 MP4s and a WebM that are
|
||||
currently just "here is a file you cannot read".
|
||||
@@ -84,8 +84,14 @@ student, and nearly free once the sync job exists.
|
||||
- **Collaborative text editor contents.** Confirmed unavailable:
|
||||
`GET /api/v3/collaborative-text-editor/{parentType}/{parentId}` returns a URL
|
||||
to the editor, never the document text.
|
||||
- **OCR.** Unnecessary — images are returned inline and Claude reads them
|
||||
directly.
|
||||
- **OCR — partially wrong, revised.** For *reading*, it remains unnecessary:
|
||||
images go to Claude inline and it reads them. For *indexing* it is a real gap.
|
||||
Measured after building the indexer: **3 of 4 sampled course PDFs have no
|
||||
embedded fonts at all** — they are scans, so extraction legitimately yields
|
||||
nothing and they are unsearchable by content. `extract.ts` now reports these
|
||||
as image-only rather than as an empty result. Making them searchable would
|
||||
need OCR (or page rasterisation plus a vision pass), which is the largest
|
||||
remaining gap in search coverage.
|
||||
- **Write tools** (submitting homework, marking tasks done). Technically easy,
|
||||
but this forfeits the property that makes an internet-facing endpoint
|
||||
acceptable: that a leaked token cannot act as the user. A deliberate, separate
|
||||
|
||||
Reference in New Issue
Block a user