Files
Schulcloud-MCP/docs/API.md
MechaCat02 a8badc2fd1 Read submitted text and teacher feedback from the homework page
You were right that this data never reaches the browser as an API call.
The legacy front end calls the Feathers API server-side for
submission.comment, submission.grade and submission.gradeComment and
renders them into GET /homework/{taskId}. That Feathers API is not
exposed publicly — /api/v1/* 404s — so the rendered page is the only way
to reach these fields from outside.

core/homework-page.ts parses it, hooked on the data-testid attributes
the project's own e2e tests use rather than incidental markup. The page
authenticates by jwt *cookie*; an Authorization header is ignored and
redirects to the identity provider. Every field is optional and parse
failures return undefined, so a markup change degrades to "not found"
and cannot break get_task. The wording distinguishes the two: absent
feedback is reported as not found, never as none given.

Measured on one course: 4 of 7 graded submissions carry feedback no API
call can return — "vollständig und nachvollziehbar", "Feedback siehe
Zettel", and so on.

This exposed a bug in a shared utility: htmlToText decoded only six
entities, so any named entity passed through raw. German content makes
that routine — "vollständig" would have reached the model verbatim
from boards and task descriptions too, not just here. It now decodes
named, decimal and hex references in one pass, so ä stays
literal instead of decoding twice, and leaves unknown names alone rather
than mangling them.

Also fixes a documented-recovery bug found while restoring the session:
`docker compose restart` does not re-read env_file, so it silently kept
serving the dead token. `up -d` is correct and the docs said the wrong
thing.

78 tests, 38/38 smoke; verified end to end through Claude Code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 13:25:35 +02:00

188 lines
9.5 KiB
Markdown

# The Schulcloud API, as verified against this instance
Everything here was confirmed against `https://schulcloud-thueringen.de` with a
real student account on 2026-09-11, not inferred from source. Where upstream
source and live behaviour disagreed, live behaviour won.
## Two services, one origin
| Service | Source repo | Base path | Self-documenting at |
|---|---|---|---|
| Main server (NestJS) | [`schulcloud-server`](https://github.com/hpi-schul-cloud/schulcloud-server) | `/api/v3/` | `/api/v3/docs`, `/api/v3/docs-json` |
| Files storage | [`file-storage`](https://github.com/hpi-schul-cloud/file-storage) | `/api/v3/file/` | `/api/v3/file/docs`, `/api/v3/file/docs-json` |
Both accept the same bearer token. The files service was split out of
`schulcloud-server` into its own repository, which is why no `file` paths
appear in the main `docs-json` — a detail that will send you in circles if you
only read the main spec. Fetch both:
```bash
curl -s "$TSC_URL/api/v3/docs-json" -o docs-v3.json # 212 paths
curl -s "$TSC_URL/api/v3/file/docs-json" -o docs-file.json # 26 paths
```
These are the authoritative reference for *this* instance's deployed version.
Prefer them over the GitHub sources, which track `main` and may be ahead.
## Which repositories matter
The `hpi-schul-cloud` org has ~100 repos, most archived or superseded. The live
ones relevant here:
- **`schulcloud-server`** — the API. Read `apps/server/src/modules/<module>/api/`
for controllers and DTOs.
- **`file-storage`** — the files service, extracted from the above.
`src/modules/files-storage/api/controller/files-storage.controller.ts` is the
whole surface.
- **`nuxt-client`** — the current web front end. Useful for seeing which API
calls the real UI makes in which order.
- **`schulcloud-client`** — the *legacy* Handlebars front end. Still receives
commits, but it is not where new features land.
Superseded/archived and worth ignoring: `authorization-service`,
`schulcloud-editor`, `nexboard-api-js`, `end-to-end-tests`, `docker-compose`,
`H5P-Nodejs-library`, `shd-client`.
Note the naming: `/api/v1` is the old Feathers surface. On this instance
`/api/v1/docs` 404s, and the v3 NestJS API covers everything this server needs.
## Content model
```
Course ─┬─ column board ─── column ─── card ─── element ─┬─ richText
│ ├─ file ──── fileRecord(s)
│ ├─ link
│ └─ …
├─ lesson (Thema) ─── contents[] + materials[]
└─ task (Aufgabe) ─── description + fileRecord(s)
```
On the account this was built against: 26 courses holding 30 column boards, 18
lessons, 42 tasks and 175 files. **Column boards hold the great majority of
current material**; lessons are the older format.
## Endpoints this server uses
| Purpose | Call |
|---|---|
| Identity, school id, permissions | `GET /api/v3/me` |
| Courses | `GET /api/v3/courses?skip&limit` |
| One course's contents | `GET /api/v3/course-rooms/{courseId}/board` |
| Dashboard tiles | `GET /api/v3/dashboard` |
| Tasks | `GET /api/v3/tasks`, `GET /api/v3/tasks/finished` |
| Lesson body | `GET /api/v3/lessons/{lessonId}` |
| Lesson's tasks | `GET /api/v3/lessons/{lessonId}/tasks` |
| Board structure | `GET /api/v3/boards/{boardId}` |
| What a board belongs to | `GET /api/v3/boards/{boardId}/context` |
| Card bodies | `GET /api/v3/cards?ids=<id>&ids=<id>` |
| Files of an entity | `GET /api/v3/file/list/{storageLocation}/{storageLocationId}/{parentType}/{parentId}` |
| One file's metadata | `GET /api/v3/file/{fileRecordId}` |
| File bytes | `GET /api/v3/file/download/{fileRecordId}/{fileName}` |
| News | `GET /api/v3/news` |
| Instance settings (no auth) | `GET /api/v3/config/public` |
| Remaining idle budget | `POST /api/v3/authentication/refresh-session``{expiresInSeconds}` |
### Gotchas that cost real time
**`course-rooms`, not `courses`, for course contents.** `GET /api/v3/courses/{id}`
does not exist. The route that returns a course's lessons/tasks/boards is
`GET /api/v3/course-rooms/{roomId}/board`, and its `:roomId` is the *course* id.
Nothing in the naming suggests this.
**`/api/v3/rooms` is a different feature.** "Rooms" are the newer standalone
collaboration spaces, unrelated to courses. On this instance the account has
none, so `GET /api/v3/rooms` returns `{"data":[]}` — which reads like a broken
endpoint but is simply an empty feature.
**`limit` maxima are enforced and mis-documented.** The OpenAPI schema says
`maximum: 99`; the runtime validator rejects anything `> 100`. Page at 99 to
satisfy both. Asking for 200 returns a `400 API_VALIDATION_ERROR`, not a
truncated list.
**There is no `GET /tasks/{id}`.** Single-task detail has to be assembled: the
list endpoints give metadata but *omit `description`*, which appears only on the
course page's task element. `get_task` does this join.
**Board files need three calls.** A `file` element's `content` carries only
`{caption, alternativeText}` — no file id. The bytes are found by listing
files-storage with `parentType: 'boardnodes'` and the **element** id as
`parentId`. This is the single least discoverable part of the API, and applies
equally to `fileFolder` and `drawing` elements.
**`GET /cards?ids=` accepts at most 20 ids.** Above that the request fails with
`400 "each value in ids must be a mongodb id"` — blaming the ids when the real
problem is how many there are. Express/NestJS parse the query string with `qs`,
whose default `arrayLimit` is 20; past it, repeated params become an object
keyed `"0"`, `"1"`, … and `@IsMongoId({ each: true })` then rejects every value.
Nothing in the controller or its DTO says so: the limit lives in the query
parser underneath them. Verified live — 20 ids return 200, 21 return 400 with
identical ids. A board with more than 20 cards is therefore unreadable in one
request; `getCards` chunks at 20.
**Submissions are nearly invisible.** The only route is
`GET /api/v3/submissions/status/task/{taskId}` — there is no `GET /submissions`
and no fetch-by-id, so a task id is the only way to reach a submission. The
response carries `{id, submitters, isSubmitted, isGraded, grade,
submittingCourseGroupName}` and **nothing else**: no submitted text, no grade
comment, no graded-at. Those lived on the legacy Feathers API, and `/api/v1` is
not served on this instance (404 across the board), so they are simply
unavailable. Submitted *files* are reachable through files-storage with
`parentType: 'submissions'`.
**Submitted text and written feedback exist only in the rendered web page.**
The legacy front end calls the Feathers API server-side for
`submission.comment`, `submission.grade` and `submission.gradeComment` and
renders them into `GET /homework/{taskId}`. That Feathers API is not exposed
publicly (`/api/v1/*` 404s), so the page is the only way to reach these from
outside. `core/homework-page.ts` parses it, hooked on the `data-testid`
attributes the project's own e2e tests use. Note the page authenticates with the
`jwt` **cookie** — an `Authorization` header is ignored and redirects to the
identity provider. Measured: 4 of 7 graded submissions in one course carried
feedback no API call can return.
**files-storage ignores `parentType` when listing.** Asking for
`.../gradings/{submissionId}` returns the files parented to that id whatever
their type — the records come back saying `parentType: "submissions"`. The path
segment appears to serve authorisation, not filtering, so filter on each
record's own `parentType` or a student's own upload will be reported back as
teacher feedback.
**`storageLocationId` is the school id** (from `/me`), with
`storageLocation: 'school'`, for every parent type in normal use.
**Lesson ids come back as buffers.** `GET /api/v3/lessons/{id}` returns nested
ids as `{buffer:{type:'Buffer',data:[...]}}` rather than hex strings — a leak
from the legacy Mongo serialisation. `normalizeObjectId` in `src/render.ts`
converts them.
**The JWT's `exp` is not the session lifetime, and the source misleads here.**
Both the current and legacy whitelist implementations re-set a Valkey TTL on
every authenticated request, which reads as a sliding window. The live instance
does not behave that way: a session ends ~2 h after **login**, and successful
reads in between do not extend it (measured — see `docs/AUTH.md`). This is the
clearest case in this API of live behaviour diverging from upstream source.
**`GET /api/v3/config/public` is unauthenticated and useful.** 78 keys of
instance configuration, including the session timeouts and feature flags. Handy
for checking deployed settings without a token.
**`Content-Disposition` on downloads is malformed.** It comes back as
`attachment;; filename="…"` — note the doubled semicolon — and the filename is
percent-encoded inside the quotes. Parse defensively.
### Content element types
From `ContentElementType` in `schulcloud-server`, all seen live except where
noted: `richText`, `file`, `fileFolder`, `link`, `drawing`,
`collaborativeTextEditor`, `externalTool`, `videoConference`, `h5p`, `deleted`.
Collaborative text editor contents are **not** retrievable through the API —
`GET /api/v3/collaborative-text-editor/{parentType}/{parentId}` returns a URL to
the Etherpad-style editor, not the document text.
## Re-verifying after an upstream release
`npm run probe` re-checks every assumption above against the live instance and
prints what it finds — both token clocks included: days until hard expiry, and
seconds of idle budget remaining.