You were right that this data never reaches the browser as an API call.
The legacy front end calls the Feathers API server-side for
submission.comment, submission.grade and submission.gradeComment and
renders them into GET /homework/{taskId}. That Feathers API is not
exposed publicly — /api/v1/* 404s — so the rendered page is the only way
to reach these fields from outside.
core/homework-page.ts parses it, hooked on the data-testid attributes
the project's own e2e tests use rather than incidental markup. The page
authenticates by jwt *cookie*; an Authorization header is ignored and
redirects to the identity provider. Every field is optional and parse
failures return undefined, so a markup change degrades to "not found"
and cannot break get_task. The wording distinguishes the two: absent
feedback is reported as not found, never as none given.
Measured on one course: 4 of 7 graded submissions carry feedback no API
call can return — "vollständig und nachvollziehbar", "Feedback siehe
Zettel", and so on.
This exposed a bug in a shared utility: htmlToText decoded only six
entities, so any named entity passed through raw. German content makes
that routine — "vollständig" would have reached the model verbatim
from boards and task descriptions too, not just here. It now decodes
named, decimal and hex references in one pass, so ä stays
literal instead of decoding twice, and leaves unknown names alone rather
than mangling them.
Also fixes a documented-recovery bug found while restoring the session:
`docker compose restart` does not re-read env_file, so it silently kept
serving the dead token. `up -d` is correct and the docs said the wrong
thing.
78 tests, 38/38 smoke; verified end to end through Claude Code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
6.1 KiB
Extensions: built and possible
Items 1–3 are now implemented — see docs/CLI.md and the store/,
indexer/ and cli/ modules. What remains below is the rationale (worth
keeping) and the items still open.
Decisions already taken
| Decision | Choice | Rejected |
|---|---|---|
| Endpoint auth | Static bearer token (MCP_AUTH_TOKEN, constant-time check) |
OAuth 2.1; secret URL path |
| State | Stateless — every tool call hits the API live | Postgres cache; cache + full-text search |
| File delivery | Inline extracted text; images inline | Raw base64 only |
| If persistence is added | The Pi's existing PostgreSQL, with its own database and user | A dedicated container; SQLite |
Two of the rejected options need no revisiting. OAuth 2.1 is disproportionate
machinery for a single-user endpoint a bearer token already protects. "Raw bytes
as well as extracted text" got built anyway — download_file takes raw: true.
1. Full-text search over file contents — BUILT
The gap: search covers course/board/card/lesson/task titles and text, and
file names — never file contents. The 93 PDFs in this account are opaque to
it. "Where does it explain the Caesar cipher?" only hits if those words are in a
filename.
The cost: one search walks ~26 course pages, ~30 board skeletons plus card
fetches, and ~150 per-element file listings — roughly 270 upstream requests,
several seconds, every time.
Design sketch
- Postgres
tsvector+ GIN. Use thegermandictionary, so Verschlüsselung matches Verschlüsselungsverfahren; the defaultenglishconfig stems these wrongly. - Add
pg_trgmalongside. German compounds defeat stemming — Datenschutzgrundverordnung will not match Datenschutz — and trigram similarity also absorbs typos. - Index: titles and text of courses, boards, cards, lessons, tasks; file names; and extracted file text.
- Later, optionally: embeddings in
pgvector, so "the thing about encryption" finds Verschlüsselung without sharing a word. Needs an embedding model in the loop, so it is a separate step, not part of this.
2. Cache, with bypass — BUILT (search fresh=true, refresh_index)
The split that makes staleness tolerable — index is discovery, live API is detail. Search the index to find where something is; always re-fetch it to read it. A stale index then costs a missed or spurious hit, never wrong content. That is a far better failure mode than a cache that can serve last week's homework list as current.
fresh: trueonsearch(re-crawl the relevant course first) and on the detail tools.- State freshness in every cached result ("indexed 2h ago"), regardless of the flag — a bypass only helps if the agent knows it needs one.
- Extracted file text is cacheable indefinitely. A
fileRecordid maps to immutable bytes; an edit produces a new record. PDF parsing happens once, ever. Cheapest large win here, and free of staleness risk. - Sync: nightly full crawl (~270 requests, trivial) plus on-demand per course.
Consequence to weigh: coursework text would then live on the Pi at rest, outside Schulcloud. It does not weaken the read-only property, but it is new — consider disk encryption and whether it lands in a backup.
3. "What's new since …" — BUILT (what_changed)
Falls out of having an index with history: diff successive crawls to surface new boards, cards, files and tasks. Impossible today at any speed — the API has no changed-since filter anywhere. Arguably the most useful item on this list for a student, and nearly free once the sync job exists.
4. Remaining gaps, measured
Scanned the instance's 98 GET routes against what the tools cover. Almost everything student-facing is now covered; what is left, with live checks:
| Area | State on this account |
|---|---|
GET /groups/class |
3 real classes — the only uncovered endpoint with data |
GET /rooms |
empty — the feature is unused here |
GET /media-boards/me |
exists, no content |
GET /course-info |
403 — teacher/admin only |
GET /alert |
empty |
tools, oauth2, school, registrations, systems, user-login-migrations |
admin/auth plumbing, not student-facing |
Genuinely unavailable, not merely uncovered:
- Collaborative text editor contents — the endpoint returns an editor URL, never the document text.
- Calendar — a separate service, absent from the v3 document.
- Image-only PDFs — 22 of 255 indexed files are scans with no text layer, so their contents cannot be searched without OCR.
- Numeric grades — the API's
gradewas null on every graded submission here, so that path stays unverified against real data.
5. Still open
- Video/audio transcription — this account has 5 MP4s and a WebM that are currently just "here is a file you cannot read".
- MCP resources and prompts — expose courses/boards as attachable resources; canned prompts such as "summarise this week's homework".
- Calendar — a separate
schulcloud-calendarservice, not part of the v3 API mapped inAPI.md.
5. Ruled out
- Collaborative text editor contents. Confirmed unavailable:
GET /api/v3/collaborative-text-editor/{parentType}/{parentId}returns a URL to the editor, never the document text. - OCR — partially wrong, revised. For reading, it remains unnecessary:
images go to Claude inline and it reads them. For indexing it is a real gap.
Measured after building the indexer: 3 of 4 sampled course PDFs have no
embedded fonts at all — they are scans, so extraction legitimately yields
nothing and they are unsearchable by content.
extract.tsnow reports these as image-only rather than as an empty result. Making them searchable would need OCR (or page rasterisation plus a vision pass), which is the largest remaining gap in search coverage. - Write tools (submitting homework, marking tasks done). Technically easy,
but this forfeits the property that makes an internet-facing endpoint
acceptable: that a leaked token cannot act as the user. A deliberate, separate
decision — see the invariant in
CLAUDE.md.