Files
Schulcloud-MCP/CLAUDE.md
MechaCat02 5ae2210459 Close the gaps an audit of courses, tasks, files and grades turned up
Every area — courses, rooms, boards, topics, tasks, files, quizzes, teams,
groups, submissions, grades — was checked for data the instance has and the
tools did not show.

Grades and feedback. A teacher's /homework page is a different page from a
student's: grade and comment live in the grading form, one block per
submission, so a teacher account reported every graded submission as having
neither. parseTeacherGrading reads the form, and list_submissions can now
include the written feedback and who handed the work in.

Names. /api/v1 is partly served: courses, users and classes survive in the
deployment's ingress table, and users/{id} is the only route from an id to a
name. Submitters, file creators and course teachers resolve through it, and
degrade to "not visible to this account" where a student may not read them.

Courses, rooms and classes. get_course adds the description, teachers,
member count and weekly timetable from /api/v1/courses. list_classes is new.
get_room reports what the account may do — allowedOperations is an object of
booleans, not the list it was typed as — and applicants and invitation links
where it may manage them.

Board and topic content. Link descriptions, image alt text, drawing and
video-conference titles, the ids behind external tools and H5P content (the
only thing resembling a quiz), and what a deleted element used to be. Topic
Etherpad pads are read like board pads, and htmlToText keeps table columns
apart and drops template indentation.

Files. A scan with no text layer falls back to the preview endpoint, whose
width and outputFormat are undocumented enums, so Claude gets a picture of
the page; list_files reports counts and sizes. Teams stay documented as
unreadable at any API version; their files come later.

What the crawl missed. Tasks attached to topics (18 of 60 on the live
account), each course's own file area, and — behind INDEX_PERSONAL_FILES —
personal files and submissions with their grade comments, so search and
what_changed cover grading. A submission hit points at get_task.

The local instance's preview profile gets an ImageMagick policy that allows
the coders its 7.1.2 build needs; the image's own denies them all.

110 tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 20:19:16 +02:00

262 lines
15 KiB
Markdown

# CLAUDE.md
Guidance for Claude Code when working in this repository.
## What this is
Read-only access to a Schulcloud (HPI Schul-Cloud / Schulcloud-Verbund-Software)
account: courses, column boards, lessons, tasks, files with text extraction, and
a Postgres-backed full-text index. TypeScript, Node 22+,
`@modelcontextprotocol/sdk`.
Three entry points over one core:
- `src/bin/http.ts` — Streamable HTTP + `/api`, the deployed form, behind Caddy on a Pi.
- `src/bin/stdio.ts` — stdio, for local Claude Code / Desktop use.
- `src/bin/cli.ts` — the `schulcloud` CLI, which talks to the HTTP server, never
to Schulcloud.
## Commands
```bash
npm run build # tsc → dist/ (also copies store/migrations/*.sql)
npm run dev # watch mode, runs src/ directly via type stripping
npm test # unit tests (node:test), no network
npm run typecheck
npm run probe # verify token + API assumptions against the LIVE instance
npm run smoke # full end-to-end: real server + real MCP client + real data
npm run keepalive-status # is the deployed container holding its session?
npm run session-diagnose # ~2.5h: measure what actually ends the session
```
`docker-compose.override.yml` is local-only and publishes the server on
`127.0.0.1:8080` and Postgres on `127.0.0.1:55432`; see `docs/LOCAL.md`.
`probe` and `smoke` hit the live Schulcloud and need a valid `.env`. Both are
read-only with respect to Schulcloud. Run `smoke` after touching `src/core/`,
`src/mcp/` or `src/http/` — the unit tests cover only pure functions.
Run smoke **both ways**: with `DATABASE_URL` set (34 checks, index-backed) and
without (32 checks, live-only). The degradation path is a supported mode, not a
fallback nobody exercises.
Store tests need a database and skip without one:
`TEST_DATABASE_URL=postgresql://… npm test`. They use a real Postgres on
purpose — the generation/diff semantics are entirely SQL, so a mock would test
nothing. **They `TRUNCATE`**, and refuse to run unless the database name
contains "test"; that guard exists because pointing them at the dev database
once put fixtures into real data.
## Architecture
```
bin/{http,stdio}.ts ─┬─ mcp/server.ts ── mcp/tools/*
└─ http/{server,api,auth}.ts /mcp and /api
bin/cli.ts ──HTTP──────────┘ cli/{config,client,sync}.ts
services.ts (process-wide: client, Store, Indexer)
indexer/indexer.ts ── store/store.ts ── Postgres
core/{client,board,crawl,extract,text,paths,types}
```
- **`core/`** knows nothing of MCP, HTTP or the CLI.
- `client.ts` — every upstream call; `GET`-only except `extendSession`.
- `board.ts` — a column board needs three kinds of call to reconstruct.
- **`crawl.ts`** — the one traversal. Search, the indexer, the what-changed
diff and the file mirror all need it; keep it here, not in a tool.
- `paths.ts` — the security boundary for mirrored filenames. See Invariants.
- **`store/`** — crawl generations, identity diffs, `german` + `pg_trgm` FTS.
`Store.open` returns `undefined` when Postgres is down; callers degrade.
- **`indexer/`** — crawl → persist → mirror bytes → extract text → index.
Coalesces concurrent refreshes; enforces a minimum interval. The crawl walks
topic-attached tasks too, which the course page does not list: without that
they are unsearchable and their grades invisible. `INDEX_PERSONAL_FILES`
additionally indexes personal files and submitted/returned work, including
grade comments — that is what makes "what got graded this week" answerable,
at roughly three extra requests per task.
- **`mcp/tools/*.ts`** — tool descriptions are prompts: they are how Claude picks
a tool, so they carry the German domain terms (Kurse, Themen, Aufgaben) and say
when *not* to use the tool.
- **`context.ts`** — per-session state. Only `/me` is cached, because the school
id is on every files-storage path and cannot change for a token.
## Invariants
**Everything that touches Schulcloud is read-only.** Every client method is a
`GET` except `extendSession` (the keepalive's `refresh-session` call, which
touches only our own session and is not exposed as a tool, so no model-driven
call can be a POST). `api_get` rejects non-`/api/` paths and anything carrying a
scheme or host. `refresh_index` and `POST /api/refresh` write only to the Pi's
own index and mirror — every upstream call they make is still a GET.
**Filenames from Schulcloud are untrusted paths.** Course titles, card titles
and filenames are all user-supplied upstream, and both the server's mirror and
the CLI's sync turn them into filesystem paths. Everything goes through
`core/paths.ts`: `safeComponent` reduces one string to one safe component, and
`resolveWithin` refuses anything that escapes the root. Do not bypass them with
`path.join`, and keep the property that no `..` survives anywhere in a
component — it is what makes the invariant checkable. The endpoint
is internet-facing by necessity, so "a leaked token cannot act as the user" is
the property that makes that acceptable. Do not add a write tool without the
user explicitly asking for one and understanding this.
**Never log or echo secrets.** `TSC_JWT_COOKIE` grants full read access to the
account; `MCP_AUTH_TOKEN` guards the endpoint. Neither belongs in
logs, error messages, or tool output. `.env` is git-ignored — keep it that way.
**Live behaviour beats upstream source.** The clones in `vendor/` track `main`
and may be ahead of what is deployed. When they disagree with the instance, the
instance is right. `docs/API.md` records which is which.
## API gotchas
These cost real time to discover; `docs/API.md` has the full list with evidence.
- Course contents are at `GET /api/v3/course-rooms/{courseId}/board`. There is
no `GET /api/v3/courses/{id}`, and `:roomId` there is the *course* id.
- **Rooms ("Räume") are a separate space from courses, and the UI's naming is a
trap**: the sidebar's *Kurse* entry links to `/rooms/courses-overview` and
lists courses; *Räume* links to `/rooms` and lists rooms. A `/rooms/...` url
says nothing about which. `list_rooms`/`get_room` cover the latter; a room
holds boards only — no lessons, no tasks. Empty is normal and is also what a
revoked membership looks like.
- Room boards report `isVisible`, which the course-page projection does not, so
a room's drafts can be named as drafts instead of being tried and 403ing.
- `limit` is rejected above 100 though the spec says 99. Page at 99; the client
clamps and `listAllCourses` pages for you.
- There is no `GET /tasks/{id}`, and the task lists omit `description` — it
only exists on the course page's task element. `get_task` does that join.
- **`GET /cards?ids=` takes at most 20 ids** (the `qs` `arrayLimit` default), and
fails above that with a validation error that blames the ids rather than their
number. `MAX_IDS_PER_QUERY` in `core/client.ts`. Any board over 20 cards is
affected, which is common.
- **Never swallow a per-item crawl error.** Board failures used to be caught and
dropped, so the index lost whole boards while the crawl reported success —
which is how the 20-id limit went unnoticed. They go into `Snapshot.failures`.
- **`GET /lessons/{id}/tasks` is a bare array whose items carry no id.** Not the
`{data,total}` envelope, and `LessonLinkedTaskResponse` has no id field at
all. A topic-attached task is thus unidentifiable from the API and invisible
in both task lists once past due — 18 of 60 tasks on the real account.
`core/lesson-page.ts` scrapes the ids off the legacy topic page.
- **Collaborative text editor (Etherpad) contents are reachable, in two hops.**
`GET /api/v3/collaborative-text-editor/content-element/{id}` returns the pad
url *and* sets an Etherpad `sessionID` cookie; `/etherpad/p/{id}/export/txt`
then returns the text. No Etherpad API key needed. `core/etherpad.ts` checks
the url's host before sending the cookie to it.
- **A draft board is listed on the course page but 403s when opened.** Say "not
published yet", not "no access".
- **File records are mutable**: `PATCH /file/rename/{id}` keeps the id and size,
so the store's digest has to include the name.
- **Submissions: only `GET /submissions/status/task/{taskId}` exists.** No list,
no fetch-by-id, and the payload has no submitted text, grade comment or
graded-at. Don't imply absent feedback means none was given.
- **`/api/v1` is partly served, and it is production surface.** Exactly three
legacy routes survive in the deployment's own ingress table
(`dof_app_deploy/ansible/group_vars/all/x_ingress.yml`): **`/api/v1/courses`,
`/api/v1/users`, `/api/v1/classes`**. Everything else under `/api/v1` is
unrouted and 404s. They matter because v3 dropped things they still carry:
`courses` has the description, `teacherIds`, `userIds` and `times` (the weekly
timetable), and `users/{id}` is the **only** way to turn a user id into a name
— submission `submitters`, file `creatorId` and course `teacherIds` are
otherwise unreadable. Permission is per-account: a teacher may read their
students, a student may read only themselves, so name resolution must degrade
to "not visible to this account" rather than printing a bare id.
- **The teacher's homework page is a different page from the student's.** Its
tabs are `extended` and `submissions`, not `submission` and `feedback`, and
the grade lives in the grading *form* (`name="grade"`, `name="gradeComment"`,
one block per `submissionId`) rather than in rendered prose. The student
parser finds nothing on it, which is why a teacher account reported every
graded submission as "neither a percentage nor feedback was found" while the
data was plainly there. `parseTeacherGrading` handles that side.
- **A grade is a percentage (`Number` 0-100) or absent; there is no text grade.**
Teachers commonly grade with `gradeComment` alone, so "graded by feedback" is
a complete answer. `formatGradeState` in `mcp/tools/submissions.ts` owns that
wording — don't reintroduce "no numeric grade recorded", which reads as a
fault.
- **Submitted text and grade comments are scraped, not fetched.** No API
exposes them; the legacy page `GET /homework/{taskId}` renders them, and it
authenticates by `jwt` **cookie**, not bearer. `core/homework-page.ts` parses
it on `data-testid` hooks and every field is optional — a markup change must
degrade to "not found", never break `get_task`.
- **files-storage listing ignores the `parentType` path segment** — filter on
each record's own `parentType`, or submission files get reported as grading
files.
- **Board file elements carry no file id.** Files are found by listing
files-storage with `parentType: 'boardnodes'` and the *element* id as
`parentId`. Same for `fileFolder` and `drawing`.
- Files live in a separate service (`/api/v3/file/*`, repo `file-storage`) with
its own OpenAPI document. It is not in the main `docs-json`.
- Legacy lesson responses return ids as `{buffer:{data:[...]}}`; use
`normalizeObjectId`.
- **`updatedAt` on the course-board projection is the request time**, not a
modification time — two reads seconds apart differ. Never build change
detection on it; the store diffs crawl generations by identity instead. The
dedicated endpoints (`/boards/{id}`, `/cards`, file records) are stable.
- Many course PDFs are **image-only scans with no text layer** (3 of 4 sampled),
so extraction legitimately yields nothing. `extract.ts` detects this and says
so; do not "fix" it by retrying. `download_file` then falls back to
`GET /file/preview/...`, which renders the page as a picture Claude can read —
the answer for a scan, though it still leaves the file unsearchable.
- **The preview endpoint has two enums, and both 400 without saying so.**
`width` accepts only **50, 150 or 500** — a number outside that set is a
validation error naming the value but not the permitted set. `outputFormat`
accepts only **`image/webp`**; omitting it is worse than wrong, because the
preview is then rendered in the *source* format and a PDF comes back as a
PDF. The response also labels itself `webp` rather than `image/webp`, so the
content type has to be normalised before anything will treat it as an image.
- **A room's `allowedOperations` is an object, not a list.** Every operation is
present with a boolean; `false` means denied. Typing it as `string[]`
type-checks and throws `.some is not a function` the moment anything reads it.
- **Schulcloud has no quiz of its own.** There is no quiz module or endpoint
upstream: interactive exercises are H5P elements, whose `contentId` is the
only handle onto the content, or external (LTI) tools behind
`contextExternalToolId`. Say that rather than looking for a quiz API.
- **Teams cannot be read at any version.** v3 exposes only
`GET /team/{teamId}/news`; upstream `main`'s teams controller is write-only
(`POST :teamId/create-room`). `/teams` is the legacy client's HTML page, not
an API. This one is genuinely unavailable, not merely uncovered.
- **`exp` (30 days) is not the session lifetime.** The binding limit is a Valkey
whitelist entry with a `JWT_TIMEOUT_SECONDS` TTL (7200s; live value at
`GET /api/v3/config/public`) that every authenticated request re-sets.
`src/keepalive.ts` holds it open — don't remove it.
- **A Schulportal tab left open revokes our token.** The `jwt` cookie *is* the
browser's session token, same `jti`. The front end runs a client-side timer
(reset only on route change, never from the server TTL) and calls
`/logout?auto-logout=true` ~2h after login, which issues `POST /api/v3/logout`
and deletes the shared key. No keepalive can prevent it; the fix is to close
the tab. This produced two false conclusions before being found — if a token
dies ~2h after login, suspect an open tab first. `docs/AUTH.md` has the chain.
## Conventions
- Imports use `.ts` extensions; `rewriteRelativeImportExtensions` makes `tsc`
emit `.js`. This lets `node --watch src/bin/http.ts` run the tree directly.
- **No TypeScript parameter properties** (`constructor(private readonly x: T)`).
Node's type stripping rejects them, which breaks `npm run dev` and `npm test`.
Declare the field and assign it in the constructor body instead.
- Tabs for indentation, single quotes, trailing commas.
- Comments explain *why* — an API quirk, a security property, a trade-off — not
what the line does. Several such comments encode findings that are expensive
to rediscover; do not strip them.
- Tool failures return `isError: true` with an actionable message via
`mcp/tools/result.ts`. `toToolError` separates 401 (token expired — the user must
act) from 403 (no access) from 404 (bad id) deliberately; keep that split.
## Adding a tool
1. Add the client method in `core/client.ts` (`GET` only).
2. Register the tool in the relevant `mcp/tools/*.ts`, with a description that
says when to use it *and when not to*.
3. Format output as Markdown, keeping ids visible for follow-up calls.
4. If it reads the index, handle `context.store === undefined` with a message
saying what is unavailable and what still works.
5. Add a check to `scripts/smoke.mjs` and run `npm run smoke` both ways.
## Environment
`.env` holds `TSC_URL`, `TSC_JWT_COOKIE`, `MCP_AUTH_TOKEN`. See `.env.example`
for the full set and `docs/AUTH.md` for refreshing the JWT. `npm run probe`
reports both clocks: days until hard expiry and seconds of idle budget left.