Files
Schulcloud-MCP/README.md
MechaCat02 bfccb3f343 Pull the server image on the Pi instead of building it
The Pi now runs registry.mc02.dev/schulcloud-mcp. Its compose file `!reset`s
the build section docker-compose.yml declares, so no `docker compose` command
there can fall back to building (Compose 2.24 or later). SCHULCLOUD_MCP_TAG
pins a commit's image; unset, the Pi follows `latest`.

`npm run publish-image` publishes from a development machine. It builds only a
clean working tree, so a commit tag names exactly that commit's code, for amd64
and arm64, tagged `latest` and with the short commit id. On an x86 machine
arm64 builds under QEMU, which the script checks for and explains rather than
registering unasked.

PI.md logs in to the registry, pulls and starts instead of building, and
updates by publishing and pulling. Rehearsed in a scratch compose project: the
image came from the registry, `up -d --build` built nothing, and the server
came up healthy with its database and no published ports.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 21:26:44 +02:00

206 lines
10 KiB
Markdown

# schulcloud-mcp
Read-only access to a [Schulcloud](https://github.com/hpi-schul-cloud) account —
courses, boards, lessons, tasks and files — for **Claude**, via MCP, and for
**you**, via a CLI that mirrors your coursework to disk.
Both are front ends over one core library and one live Schulcloud session, kept
alive on a Pi.
Built and verified against `schulcloud-thueringen.de` with a live student
account. Everything in `docs/API.md` was confirmed against the running
instance, not inferred from the upstream source.
## What Claude can do with it
> *"What do I have due this week?"*
> *"Find the material about Verschlüsselung and explain the Caesar cipher worksheet."*
> *"Summarise the routing lesson from the LF10 course."*
Twenty-four tools, all read-only:
| | |
|---|---|
| `whoami` | account, school, roles — also a connectivity check |
| `list_courses` | all courses, with ids |
| `get_dashboard` | the tiles as pinned on the web dashboard |
| `get_course` | one course's boards, topics and tasks |
| `get_board` | a column board in full: columns, cards, text, links, files |
| `get_lesson` | a topic's text sections, materials, files and tasks |
| `list_tasks` | homework across all courses, by due date |
| `get_task` | one task: description, due date, attachments, **and your submission** |
| `list_submissions` | what you handed in, and what is still ungraded |
| `list_files` | files attached to a board element, topic, task or submission |
| `download_file` | fetch such an attachment and extract its text, or view an image |
| `fs_list` | list a folder of the file manager — Persönliche, Kurs-, Team-, Geteilte Dateien |
| `fs_tree` | everything below a file-manager folder, as a tree |
| `fs_find` | find file-manager files and folders by name |
| `fs_read` | read a file-manager file, extracted like `download_file` |
| `search` | keyword search across everything — **including the text inside PDFs and Office files** |
| `refresh_index` | re-read Schulcloud now, per course or in full |
| `what_changed` | what appeared, changed or vanished since a date |
| `index_status` | how fresh the index is |
| `list_rooms` | rooms ("Räume"), which are a separate space from courses |
| `get_room` | one room: its boards, members and what you may do there |
| `list_classes` | classes ("Klassen") with their teachers, and group membership |
| `list_news` | school and course announcements |
| `api_get` | GET-only escape hatch for uncovered API surface |
`download_file` extracts text from **PDF, DOCX, XLSX, PPTX and OpenDocument**
files and returns **images inline** for Claude to look at. Image-only PDFs —
scans with no text layer, which are common in this account — are reported as
such rather than as an empty result.
**Attach a course instead of asking for it.** Every course and room is also an
MCP resource (`schulcloud://courses/<id>`, `schulcloud://rooms/<id>`) holding
the same overview `get_course` and `get_room` return. In Claude Code, type `@`
and part of the course name.
**Two ready-made prompts**, in German because the school is:
| | |
|---|---|
| `zusammenfassung` *kurs* [*fokus*] | summarise a course or room: topics, tasks, key material, each with its source |
| `pruefungsvorbereitung` *kurs* [*thema*] [*datum*] | prepare for an exam: scope, explanations, practice questions, a study plan |
In Claude Code they run as `/mcp__schulcloud__zusammenfassung Mathe_10b`.
Claude Code splits arguments on spaces and drops extra words, so join words
with `_` (`Lineare_Funktionen`) and skip an optional argument with `-`. *kurs*
is any unambiguous part of a course or room name, or its id.
## The CLI
```bash
schulcloud login --server https://mcp.example.org --token <token>
schulcloud sync --dry-run # see what would be mirrored
schulcloud sync # mirror coursework to ~/Schulcloud
schulcloud refresh --course <id>
schulcloud fs tree /courses # browse the file manager ("Dateien")
schulcloud fs get "/courses/<course>/<folder>"
schulcloud token set # the monthly chore: hand the Pi a fresh Schulcloud token
```
It talks only to the Pi and holds no Schulcloud credential — see
[docs/CLI.md](docs/CLI.md). A fresh token can also be pasted into the server's
`/token` page; either way the server checks it with Schulcloud and swaps it in
without a restart.
## Quick start
```bash
cp .env.example .env # fill in TSC_URL and TSC_JWT_COOKIE
npm install
npm run build
npm run probe # verifies the token and API against the live instance
```
To try it on your own machine — Docker stack, Claude Code, and the CLI — follow
[docs/LOCAL.md](docs/LOCAL.md). To put it on a Pi behind Caddy and a VPS, and
connect claude.ai, follow [docs/PI.md](docs/PI.md);
[docs/DEPLOYMENT.md](docs/DEPLOYMENT.md) explains the pieces.
Getting `TSC_JWT_COOKIE` takes four clicks in DevTools and then lasts 30 days —
provided you close the Schulportal window afterwards. See
[docs/AUTH.md](docs/AUTH.md); that caveat is not optional.
## Design decisions
**Bearer token plus a keepalive.** The instance's `jwt` cookie works verbatim
as `Authorization: Bearer` — no cookie jar, no `connect.sid`. Its `exp` claim
(30 days) is only a ceiling: the real limit is a 2-hour server-side session TTL
that any request slides, so the server calls `refresh-session` every 30 minutes
(the one non-GET request here, and not exposed as a tool).
The sharp edge is subtler and cost two endurance tests to find: **the cookie you
copy is the browser's own session token**, so a Schulportal tab left open will
auto-logout after ~2 hours and revoke this server's token with it. Copy the
token in a private window and close it. See [docs/AUTH.md](docs/AUTH.md).
**claude.ai reaches it by a secret path, for now.** Its connector dialog takes
a URL and no header, so `MCP_PATH_SECRET` serves MCP at `/<secret>/mcp` without
the bearer token — never logged, redacted by the Caddy snippet, and a stopgap
until the endpoint speaks OAuth. Claude Code and the CLI keep the bearer token.
See [docs/DEPLOYMENT.md](docs/DEPLOYMENT.md).
**Read-only by construction.** Every method on the API client is a `GET`,
including `api_get`. The endpoint is internet-facing by necessity (Claude's
connectors call it from Anthropic's cloud), so the fact that a leaked token
cannot be used to *act* as the user is the main safety property. Adding one
write tool would forfeit it.
**An index, with an honest bypass.** Postgres holds crawl generations and a
`german` + `pg_trgm` full-text index over extracted file text, so search covers
the inside of PDFs rather than just their names. Every result states how fresh
the index is, and `fresh=true` bypasses it for a live read — an agent should
never be quietly misled by stale data. Without `DATABASE_URL` the server still
works: search falls back to crawling live.
**Generations, not timestamps.** Sync cursors and change detection compare crawl
generations by identity. Measured: the course-board endpoint returns *request
time* as `updatedAt`, so a timestamp cursor would report everything as changed
on every crawl — and could never detect deletions.
**Assembled, not raw.** `get_board` makes three kinds of upstream call and
stitches the results — board skeleton, card bodies, and a files-storage lookup
per file element — because a model asking "what's on this board" wants the
answer, not a traversal plan. Output is Markdown with ids preserved for
follow-up calls, not raw JSON.
Possible extensions — full-text search over file contents, a cache with a
bypass, "what's new since…" — are sketched with their trade-offs in
[docs/ROADMAP.md](docs/ROADMAP.md). None are built.
## Layout
```
src/
core/ client, types, board assembly, crawler, extraction, paths
store/ Postgres: crawl generations, diffs, full-text search
indexer/ crawl → persist → mirror bytes → extract text → index
mcp/ MCP server, tools, resources and prompts
http/ express app, bearer auth, /api for the CLI
cli/ CLI config, API client, sync engine
bin/ http, stdio and cli entry points
docs/ API findings, auth, deployment, CLI, roadmap
deploy/ Caddyfile snippet, the Pi's compose file
scripts/ probe, smoke, session diagnostics
vendor/ upstream clones, git-ignored, for reference only
```
`core/` knows nothing about MCP, HTTP or the CLI: it holds the Schulcloud client,
the traversal every feature needs, document extraction, and the path
sanitisation that both the server's mirror and the CLI's sync depend on.
## Development
```bash
npm run dev # watch mode, runs src/ directly
npm test # unit tests, no network
npm run probe # check assumptions against the live instance
npm run smoke # full end-to-end: real server, real client, real data
npm run session-diagnose # instrument what actually ends the session (~2.5h)
npm run keepalive-status # is the deployed container holding its session?
npm run publish-image # build amd64 + arm64 and push to registry.mc02.dev — the Pi only pulls
npm run typecheck
```
`npm run smoke` starts the HTTP server, connects a real MCP client over
Streamable HTTP and exercises every tool against the live account — 74 checks (76 with the index)
covering the auth gate and the secret path, the protocol handshake, every content chain, file
extraction, resources and prompts, token replacement, `api_get`'s guard rails and error handling.
## Upstream
Reference clones live in `vendor/` (git-ignored):
```bash
git clone --depth 1 --filter=blob:none https://github.com/hpi-schul-cloud/schulcloud-server.git vendor/schulcloud-server
git clone --depth 1 --filter=blob:none https://github.com/hpi-schul-cloud/file-storage.git vendor/file-storage
git clone --depth 1 --filter=blob:none https://github.com/hpi-schul-cloud/nuxt-client.git vendor/nuxt-client
```
Most of that organisation's ~100 repositories are archived or superseded; those
three are the live ones that matter. The instance's own OpenAPI documents
(`/api/v3/docs-json`, `/api/v3/file/docs-json`) are more authoritative than any
of them — see [docs/API.md](docs/API.md).