MechaCat02 a18b526267 Add Postgres store, path safety, and the crawl indexer
Store: crawl generations as the sync cursor. Diffs compare generations
on entity identity plus a content digest, never on upstream timestamps —
GET /course-rooms/{id}/board returns request time as updatedAt for most
elements, so a timestamp cursor would report every board as changed on
every crawl. Identity diffing also yields deletions, which no timestamp
scheme can. A per-course crawl carries the other courses' rows forward
so every completed generation is a complete picture and any two diff
directly; without that a partial crawl reads as a mass deletion.

FTS uses the german dictionary with weighted title/body, plus a pg_trgm
arm because stemming will not match "Datenschutz" inside
"Datenschutzgrundverordnung" and German compounds make that the common
case. file_texts is keyed by file record id and deliberately outlives
generations: records are immutable upstream, so text extracted once is
valid forever and a re-crawl of unchanged content costs nothing.

Store.open returns undefined instead of throwing when Postgres is
unreachable — the index is an accelerator, and a Pi that loses its
database should get slower, not broken.

core/paths.ts is the security boundary for the mirror. Course titles,
card titles and filenames are all user-supplied upstream, so this is
where a hostile name stops being text and becomes a path. Two bugs found
by its own tests: "///" produced "---" instead of falling back, and dot
runs survived mid-component. Now no ".." can survive anywhere, which
makes the invariant checkable rather than a claim about ordering.

Indexer coalesces concurrent refreshes onto one run and enforces a
minimum interval, since a full crawl is ~270 requests from an account
that looks like a student.

9 store tests against a real Postgres (mocks would test nothing here)
and 13 path tests; 47 total.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 21:10:53 +02:00
2026-09-11 23:52:12 +02:00
2026-09-11 23:52:12 +02:00
2026-09-11 23:52:12 +02:00
2026-09-11 23:52:12 +02:00
2026-09-11 23:52:12 +02:00

schulcloud-mcp

An MCP server that gives Claude read-only access to a Schulcloud account — courses, boards, lessons, tasks — and reads the attached files, so you can ask about your coursework instead of downloading PDFs and uploading them by hand.

Built and verified against schulcloud-thueringen.de with a live student account. Everything in docs/API.md was confirmed against the running instance, not inferred from the upstream source.

What Claude can do with it

"What do I have due this week?" "Find the material about Verschlüsselung and explain the Caesar cipher worksheet." "Summarise the routing lesson from the LF10 course."

Thirteen tools, all read-only:

whoami account, school, roles — also a connectivity check
list_courses all courses, with ids
get_dashboard the tiles as pinned on the web dashboard
get_course one course's boards, topics and tasks
get_board a column board in full: columns, cards, text, links, files
get_lesson a topic's text sections, materials, files and tasks
list_tasks homework across all courses, by due date
get_task one task: description, due date, status, attachments
list_files files attached to any entity
download_file fetch a file and extract its text, or view an image
search keyword search across courses, boards, files and tasks
list_news school and course announcements
api_get GET-only escape hatch for uncovered API surface

download_file extracts text from PDF, DOCX, XLSX, PPTX and OpenDocument files and returns images inline for Claude to look at. Verified against real files in the account: a 93-file PDF corpus, DOCX, ODT and PPTX all extract correctly.

Quick start

cp .env.example .env        # fill in TSC_URL and TSC_JWT_COOKIE
npm install
npm run build
npm run probe               # verifies the token and API against the live instance

Then either deploy it as a remote connector, or point Claude Code at dist/bin/stdio.js. Both paths are in docs/DEPLOYMENT.md.

Getting TSC_JWT_COOKIE takes four clicks in DevTools and then lasts 30 days — provided you close the Schulportal window afterwards. See docs/AUTH.md; that caveat is not optional.

Design decisions

Bearer token plus a keepalive. The instance's jwt cookie works verbatim as Authorization: Bearer — no cookie jar, no connect.sid. Its exp claim (30 days) is only a ceiling: the real limit is a 2-hour server-side session TTL that any request slides, so the server calls refresh-session every 30 minutes (the one non-GET request here, and not exposed as a tool).

The sharp edge is subtler and cost two endurance tests to find: the cookie you copy is the browser's own session token, so a Schulportal tab left open will auto-logout after ~2 hours and revoke this server's token with it. Copy the token in a private window and close it. See docs/AUTH.md.

Read-only by construction. Every method on the API client is a GET, including api_get. The endpoint is internet-facing by necessity (Claude's connectors call it from Anthropic's cloud), so the fact that a leaked token cannot be used to act as the user is the main safety property. Adding one write tool would forfeit it.

Stateless. No database, despite one being available on the host. 26 courses is not a caching problem, and a cache would introduce staleness questions that live calls simply do not have.

Assembled, not raw. get_board makes three kinds of upstream call and stitches the results — board skeleton, card bodies, and a files-storage lookup per file element — because a model asking "what's on this board" wants the answer, not a traversal plan. Output is Markdown with ids preserved for follow-up calls, not raw JSON.

Possible extensions — full-text search over file contents, a cache with a bypass, "what's new since…" — are sketched with their trade-offs in docs/ROADMAP.md. None are built.

Layout

src/
  bin/          stdio and http entry points
  schulcloud/   API client, response types, board assembly
  tools/        one module per group of MCP tools
  http/         express app, bearer auth
  extract.ts    document → text
  render.ts     formatting helpers
docs/           API findings, auth, deployment
deploy/         Caddyfile snippet
scripts/        probe (verify against live) and smoke (end-to-end)
vendor/         upstream clones, git-ignored, for reference only

Development

npm run dev        # watch mode, runs src/ directly
npm test           # unit tests, no network
npm run probe      # check assumptions against the live instance
npm run smoke      # full end-to-end: real server, real client, real data
npm run session-diagnose  # instrument what actually ends the session (~2.5h)
npm run keepalive-status  # is the deployed container holding its session?
npm run typecheck

npm run smoke starts the HTTP server, connects a real MCP client over Streamable HTTP and exercises every tool against the live account — 30 checks covering the auth gate, the protocol handshake, every content chain, file extraction, api_get's guard rails and error handling.

Upstream

Reference clones live in vendor/ (git-ignored):

git clone --depth 1 --filter=blob:none https://github.com/hpi-schul-cloud/schulcloud-server.git vendor/schulcloud-server
git clone --depth 1 --filter=blob:none https://github.com/hpi-schul-cloud/file-storage.git      vendor/file-storage
git clone --depth 1 --filter=blob:none https://github.com/hpi-schul-cloud/nuxt-client.git       vendor/nuxt-client

Most of that organisation's ~100 repositories are archived or superseded; those three are the live ones that matter. The instance's own OpenAPI documents (/api/v3/docs-json, /api/v3/file/docs-json) are more authoritative than any of them — see docs/API.md.

Description
No description provided
Readme 2.3 MiB
Languages
TypeScript 76.4%
JavaScript 19.2%
Shell 1.3%
CSS 1.3%
HTML 0.7%
Other 1.1%