MechaCat02 c79f1b120d Index-backed search, /api surfaces, image-only PDF detection
search now queries the Postgres index and states its freshness in every
result, with fresh=true bypassing it for a live crawl — the agent can
always get current data rather than being quietly misled by a stale
index. Adds refresh_index (per-course by default; a full crawl is ~270
requests), what_changed (generation diff — the API has no changed-since
filter of any kind), and index_status.

/api gives the CLI its backend behind the same bearer token as /mcp:
GET /manifest (cursor + per-file status), GET /files/:id (served from
the mirror with Range support, falling back to a live proxy for files
too large to mirror), GET /status, POST /refresh. Bytes go over plain
HTTP rather than MCP because base64 in JSON-RPC costs a third more and
buffers whole files. An unresolvable manifest cursor returns 409 rather
than silently meaning "everything is new", so a client cannot be tricked
into a full re-download.

Verified end to end against the live instance and a real Postgres:
crawl -> index -> German FTS -> manifest -> ranged download, with 401
on missing token, 400 on a malformed id, and 429 on a too-soon refresh.

Two findings worth recording. The build silently omitted the .sql
migrations from dist, which the store's graceful degradation turned into
"running without the index" rather than a crash — now copied by a build
step. And 3 of 4 sampled course PDFs have no embedded fonts at all: they
are scans, so extraction legitimately yields nothing. That is now
detected and reported as image-only with OCR named as the missing piece,
instead of an indistinguishable "0 characters". It revises the roadmap's
"OCR not needed" note, which held for reading images but not for
indexing them.

49 tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 21:16:38 +02:00
2026-09-11 23:52:12 +02:00
2026-09-11 23:52:12 +02:00
2026-09-11 23:52:12 +02:00
2026-09-11 23:52:12 +02:00
2026-09-11 23:52:12 +02:00

schulcloud-mcp

An MCP server that gives Claude read-only access to a Schulcloud account — courses, boards, lessons, tasks — and reads the attached files, so you can ask about your coursework instead of downloading PDFs and uploading them by hand.

Built and verified against schulcloud-thueringen.de with a live student account. Everything in docs/API.md was confirmed against the running instance, not inferred from the upstream source.

What Claude can do with it

"What do I have due this week?" "Find the material about Verschlüsselung and explain the Caesar cipher worksheet." "Summarise the routing lesson from the LF10 course."

Thirteen tools, all read-only:

whoami account, school, roles — also a connectivity check
list_courses all courses, with ids
get_dashboard the tiles as pinned on the web dashboard
get_course one course's boards, topics and tasks
get_board a column board in full: columns, cards, text, links, files
get_lesson a topic's text sections, materials, files and tasks
list_tasks homework across all courses, by due date
get_task one task: description, due date, status, attachments
list_files files attached to any entity
download_file fetch a file and extract its text, or view an image
search keyword search across courses, boards, files and tasks
list_news school and course announcements
api_get GET-only escape hatch for uncovered API surface

download_file extracts text from PDF, DOCX, XLSX, PPTX and OpenDocument files and returns images inline for Claude to look at. Verified against real files in the account: a 93-file PDF corpus, DOCX, ODT and PPTX all extract correctly.

Quick start

cp .env.example .env        # fill in TSC_URL and TSC_JWT_COOKIE
npm install
npm run build
npm run probe               # verifies the token and API against the live instance

Then either deploy it as a remote connector, or point Claude Code at dist/bin/stdio.js. Both paths are in docs/DEPLOYMENT.md.

Getting TSC_JWT_COOKIE takes four clicks in DevTools and then lasts 30 days — provided you close the Schulportal window afterwards. See docs/AUTH.md; that caveat is not optional.

Design decisions

Bearer token plus a keepalive. The instance's jwt cookie works verbatim as Authorization: Bearer — no cookie jar, no connect.sid. Its exp claim (30 days) is only a ceiling: the real limit is a 2-hour server-side session TTL that any request slides, so the server calls refresh-session every 30 minutes (the one non-GET request here, and not exposed as a tool).

The sharp edge is subtler and cost two endurance tests to find: the cookie you copy is the browser's own session token, so a Schulportal tab left open will auto-logout after ~2 hours and revoke this server's token with it. Copy the token in a private window and close it. See docs/AUTH.md.

Read-only by construction. Every method on the API client is a GET, including api_get. The endpoint is internet-facing by necessity (Claude's connectors call it from Anthropic's cloud), so the fact that a leaked token cannot be used to act as the user is the main safety property. Adding one write tool would forfeit it.

Stateless. No database, despite one being available on the host. 26 courses is not a caching problem, and a cache would introduce staleness questions that live calls simply do not have.

Assembled, not raw. get_board makes three kinds of upstream call and stitches the results — board skeleton, card bodies, and a files-storage lookup per file element — because a model asking "what's on this board" wants the answer, not a traversal plan. Output is Markdown with ids preserved for follow-up calls, not raw JSON.

Possible extensions — full-text search over file contents, a cache with a bypass, "what's new since…" — are sketched with their trade-offs in docs/ROADMAP.md. None are built.

Layout

src/
  bin/          stdio and http entry points
  schulcloud/   API client, response types, board assembly
  tools/        one module per group of MCP tools
  http/         express app, bearer auth
  extract.ts    document → text
  render.ts     formatting helpers
docs/           API findings, auth, deployment
deploy/         Caddyfile snippet
scripts/        probe (verify against live) and smoke (end-to-end)
vendor/         upstream clones, git-ignored, for reference only

Development

npm run dev        # watch mode, runs src/ directly
npm test           # unit tests, no network
npm run probe      # check assumptions against the live instance
npm run smoke      # full end-to-end: real server, real client, real data
npm run session-diagnose  # instrument what actually ends the session (~2.5h)
npm run keepalive-status  # is the deployed container holding its session?
npm run typecheck

npm run smoke starts the HTTP server, connects a real MCP client over Streamable HTTP and exercises every tool against the live account — 30 checks covering the auth gate, the protocol handshake, every content chain, file extraction, api_get's guard rails and error handling.

Upstream

Reference clones live in vendor/ (git-ignored):

git clone --depth 1 --filter=blob:none https://github.com/hpi-schul-cloud/schulcloud-server.git vendor/schulcloud-server
git clone --depth 1 --filter=blob:none https://github.com/hpi-schul-cloud/file-storage.git      vendor/file-storage
git clone --depth 1 --filter=blob:none https://github.com/hpi-schul-cloud/nuxt-client.git       vendor/nuxt-client

Most of that organisation's ~100 repositories are archived or superseded; those three are the live ones that matter. The instance's own OpenAPI documents (/api/v3/docs-json, /api/v3/file/docs-json) are more authoritative than any of them — see docs/API.md.

Description
No description provided
Readme 2.3 MiB
Languages
TypeScript 76.4%
JavaScript 19.2%
Shell 1.3%
CSS 1.3%
HTML 0.7%
Other 1.1%