The CLI talks only to the Pi's /api surface and holds no Schulcloud credential — only the same bearer token the Claude connector uses. That is not layering for its own sake: a Schulcloud session dies after two hours idle and a CLI process lives for seconds, so a CLI with its own token would be dead most times you reached for it. Routing through the Pi means one session, one keepalive, one monthly cookie paste. sync is a one-way mirror, which follows from the data rather than from scope-cutting: file records are immutable upstream, so there is no versioning, no conflict resolution and no merge. State is keyed by file record id with the path as derived output, so an upstream rename moves the local file instead of duplicating it — verified against the live server. Verification is size-only because the download endpoint exposes no ETag and Schulcloud publishes no hash; size still catches the failure that happens, a truncated download. Downloads land on a .part neighbour and are renamed, so an interrupted run leaves no half-file that a later run mistakes for complete. Deletions are reported but not propagated — a teacher removing a worksheet is no reason to destroy the student's copy — with --prune to opt in. what_changed now clamps to the oldest stored generation instead of refusing, and says it did: "what's new this week" is a reasonable question to ask a two-day-old index. Two build bugs caught by the checks rather than by luck: the smoke harness constructed the app without services, so the index-backed tools were never exercised; and the Docker build could not see scripts/copy-assets.mjs, so the image would have shipped without migrations and silently degraded to live-only. 67 unit tests (9 needing Postgres), smoke green both ways — 34 checks with an index, 32 without, because graceful degradation is a supported mode and not a fallback nobody runs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
167 lines
7.4 KiB
Markdown
167 lines
7.4 KiB
Markdown
# schulcloud-mcp
|
|
|
|
Read-only access to a [Schulcloud](https://github.com/hpi-schul-cloud) account —
|
|
courses, boards, lessons, tasks and files — for **Claude**, via MCP, and for
|
|
**you**, via a CLI that mirrors your coursework to disk.
|
|
|
|
Both are front ends over one core library and one live Schulcloud session, kept
|
|
alive on a Pi.
|
|
|
|
Built and verified against `schulcloud-thueringen.de` with a live student
|
|
account. Everything in `docs/API.md` was confirmed against the running
|
|
instance, not inferred from the upstream source.
|
|
|
|
## What Claude can do with it
|
|
|
|
> *"What do I have due this week?"*
|
|
> *"Find the material about Verschlüsselung and explain the Caesar cipher worksheet."*
|
|
> *"Summarise the routing lesson from the LF10 course."*
|
|
|
|
Thirteen tools, all read-only:
|
|
|
|
| | |
|
|
|---|---|
|
|
| `whoami` | account, school, roles — also a connectivity check |
|
|
| `list_courses` | all courses, with ids |
|
|
| `get_dashboard` | the tiles as pinned on the web dashboard |
|
|
| `get_course` | one course's boards, topics and tasks |
|
|
| `get_board` | a column board in full: columns, cards, text, links, files |
|
|
| `get_lesson` | a topic's text sections, materials, files and tasks |
|
|
| `list_tasks` | homework across all courses, by due date |
|
|
| `get_task` | one task: description, due date, status, attachments |
|
|
| `list_files` | files attached to any entity |
|
|
| `download_file` | fetch a file and extract its text, or view an image |
|
|
| `search` | keyword search across everything — **including the text inside PDFs and Office files** |
|
|
| `refresh_index` | re-read Schulcloud now, per course or in full |
|
|
| `what_changed` | what appeared, changed or vanished since a date |
|
|
| `index_status` | how fresh the index is |
|
|
| `list_news` | school and course announcements |
|
|
| `api_get` | GET-only escape hatch for uncovered API surface |
|
|
|
|
`download_file` extracts text from **PDF, DOCX, XLSX, PPTX and OpenDocument**
|
|
files and returns **images inline** for Claude to look at. Image-only PDFs —
|
|
scans with no text layer, which are common in this account — are reported as
|
|
such rather than as an empty result.
|
|
|
|
## The CLI
|
|
|
|
```bash
|
|
schulcloud login --server https://mcp.example.org --token <token>
|
|
schulcloud sync --dry-run # see what would be mirrored
|
|
schulcloud sync # mirror coursework to ~/Schulcloud
|
|
schulcloud refresh --course <id>
|
|
```
|
|
|
|
It talks only to the Pi and holds no Schulcloud credential — see
|
|
[docs/CLI.md](docs/CLI.md).
|
|
|
|
## Quick start
|
|
|
|
```bash
|
|
cp .env.example .env # fill in TSC_URL and TSC_JWT_COOKIE
|
|
npm install
|
|
npm run build
|
|
npm run probe # verifies the token and API against the live instance
|
|
```
|
|
|
|
Then either deploy it as a remote connector, or point Claude Code at
|
|
`dist/bin/stdio.js`. Both paths are in [docs/DEPLOYMENT.md](docs/DEPLOYMENT.md).
|
|
|
|
Getting `TSC_JWT_COOKIE` takes four clicks in DevTools and then lasts 30 days —
|
|
provided you close the Schulportal window afterwards. See
|
|
[docs/AUTH.md](docs/AUTH.md); that caveat is not optional.
|
|
|
|
## Design decisions
|
|
|
|
**Bearer token plus a keepalive.** The instance's `jwt` cookie works verbatim
|
|
as `Authorization: Bearer` — no cookie jar, no `connect.sid`. Its `exp` claim
|
|
(30 days) is only a ceiling: the real limit is a 2-hour server-side session TTL
|
|
that any request slides, so the server calls `refresh-session` every 30 minutes
|
|
(the one non-GET request here, and not exposed as a tool).
|
|
|
|
The sharp edge is subtler and cost two endurance tests to find: **the cookie you
|
|
copy is the browser's own session token**, so a Schulportal tab left open will
|
|
auto-logout after ~2 hours and revoke this server's token with it. Copy the
|
|
token in a private window and close it. See [docs/AUTH.md](docs/AUTH.md).
|
|
|
|
**Read-only by construction.** Every method on the API client is a `GET`,
|
|
including `api_get`. The endpoint is internet-facing by necessity (Claude's
|
|
connectors call it from Anthropic's cloud), so the fact that a leaked token
|
|
cannot be used to *act* as the user is the main safety property. Adding one
|
|
write tool would forfeit it.
|
|
|
|
**An index, with an honest bypass.** Postgres holds crawl generations and a
|
|
`german` + `pg_trgm` full-text index over extracted file text, so search covers
|
|
the inside of PDFs rather than just their names. Every result states how fresh
|
|
the index is, and `fresh=true` bypasses it for a live read — an agent should
|
|
never be quietly misled by stale data. Without `DATABASE_URL` the server still
|
|
works: search falls back to crawling live.
|
|
|
|
**Generations, not timestamps.** Sync cursors and change detection compare crawl
|
|
generations by identity. Measured: the course-board endpoint returns *request
|
|
time* as `updatedAt`, so a timestamp cursor would report everything as changed
|
|
on every crawl — and could never detect deletions.
|
|
|
|
**Assembled, not raw.** `get_board` makes three kinds of upstream call and
|
|
stitches the results — board skeleton, card bodies, and a files-storage lookup
|
|
per file element — because a model asking "what's on this board" wants the
|
|
answer, not a traversal plan. Output is Markdown with ids preserved for
|
|
follow-up calls, not raw JSON.
|
|
|
|
Possible extensions — full-text search over file contents, a cache with a
|
|
bypass, "what's new since…" — are sketched with their trade-offs in
|
|
[docs/ROADMAP.md](docs/ROADMAP.md). None are built.
|
|
|
|
## Layout
|
|
|
|
```
|
|
src/
|
|
core/ client, types, board assembly, crawler, extraction, paths
|
|
store/ Postgres: crawl generations, diffs, full-text search
|
|
indexer/ crawl → persist → mirror bytes → extract text → index
|
|
mcp/ MCP server and tools
|
|
http/ express app, bearer auth, /api for the CLI
|
|
cli/ CLI config, API client, sync engine
|
|
bin/ http, stdio and cli entry points
|
|
docs/ API findings, auth, deployment, CLI, roadmap
|
|
deploy/ Caddyfile snippet
|
|
scripts/ probe, smoke, session diagnostics
|
|
vendor/ upstream clones, git-ignored, for reference only
|
|
```
|
|
|
|
`core/` knows nothing about MCP, HTTP or the CLI: it holds the Schulcloud client,
|
|
the traversal every feature needs, document extraction, and the path
|
|
sanitisation that both the server's mirror and the CLI's sync depend on.
|
|
|
|
## Development
|
|
|
|
```bash
|
|
npm run dev # watch mode, runs src/ directly
|
|
npm test # unit tests, no network
|
|
npm run probe # check assumptions against the live instance
|
|
npm run smoke # full end-to-end: real server, real client, real data
|
|
npm run session-diagnose # instrument what actually ends the session (~2.5h)
|
|
npm run keepalive-status # is the deployed container holding its session?
|
|
npm run typecheck
|
|
```
|
|
|
|
`npm run smoke` starts the HTTP server, connects a real MCP client over
|
|
Streamable HTTP and exercises every tool against the live account — 30 checks
|
|
covering the auth gate, the protocol handshake, every content chain, file
|
|
extraction, `api_get`'s guard rails and error handling.
|
|
|
|
## Upstream
|
|
|
|
Reference clones live in `vendor/` (git-ignored):
|
|
|
|
```bash
|
|
git clone --depth 1 --filter=blob:none https://github.com/hpi-schul-cloud/schulcloud-server.git vendor/schulcloud-server
|
|
git clone --depth 1 --filter=blob:none https://github.com/hpi-schul-cloud/file-storage.git vendor/file-storage
|
|
git clone --depth 1 --filter=blob:none https://github.com/hpi-schul-cloud/nuxt-client.git vendor/nuxt-client
|
|
```
|
|
|
|
Most of that organisation's ~100 repositories are archived or superseded; those
|
|
three are the live ones that matter. The instance's own OpenAPI documents
|
|
(`/api/v3/docs-json`, `/api/v3/file/docs-json`) are more authoritative than any
|
|
of them — see [docs/API.md](docs/API.md).
|