Files
Schulcloud-MCP/docs/ROADMAP.md
MechaCat02 a8badc2fd1 Read submitted text and teacher feedback from the homework page
You were right that this data never reaches the browser as an API call.
The legacy front end calls the Feathers API server-side for
submission.comment, submission.grade and submission.gradeComment and
renders them into GET /homework/{taskId}. That Feathers API is not
exposed publicly — /api/v1/* 404s — so the rendered page is the only way
to reach these fields from outside.

core/homework-page.ts parses it, hooked on the data-testid attributes
the project's own e2e tests use rather than incidental markup. The page
authenticates by jwt *cookie*; an Authorization header is ignored and
redirects to the identity provider. Every field is optional and parse
failures return undefined, so a markup change degrades to "not found"
and cannot break get_task. The wording distinguishes the two: absent
feedback is reported as not found, never as none given.

Measured on one course: 4 of 7 graded submissions carry feedback no API
call can return — "vollständig und nachvollziehbar", "Feedback siehe
Zettel", and so on.

This exposed a bug in a shared utility: htmlToText decoded only six
entities, so any named entity passed through raw. German content makes
that routine — "vollständig" would have reached the model verbatim
from boards and task descriptions too, not just here. It now decodes
named, decimal and hex references in one pass, so ä stays
literal instead of decoding twice, and leaves unknown names alone rather
than mangling them.

Also fixes a documented-recovery bug found while restoring the session:
`docker compose restart` does not re-read env_file, so it silently kept
serving the dead token. `up -d` is correct and the docs said the wrong
thing.

78 tests, 38/38 smoke; verified end to end through Claude Code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 13:25:35 +02:00

123 lines
6.1 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Extensions: built and possible
Items 13 are **now implemented** — see `docs/CLI.md` and the `store/`,
`indexer/` and `cli/` modules. What remains below is the rationale (worth
keeping) and the items still open.
## Decisions already taken
| Decision | Choice | Rejected |
|---|---|---|
| Endpoint auth | Static bearer token (`MCP_AUTH_TOKEN`, constant-time check) | OAuth 2.1; secret URL path |
| State | Stateless — every tool call hits the API live | Postgres cache; cache + full-text search |
| File delivery | Inline extracted text; images inline | Raw base64 only |
| **If** persistence is added | **The Pi's existing PostgreSQL, with its own database and user** | A dedicated container; SQLite |
Two of the rejected options need no revisiting. OAuth 2.1 is disproportionate
machinery for a single-user endpoint a bearer token already protects. "Raw bytes
as well as extracted text" got built anyway — `download_file` takes `raw: true`.
## 1. Full-text search over file contents — BUILT
**The gap:** `search` covers course/board/card/lesson/task titles and text, and
file *names* — never file *contents*. The 93 PDFs in this account are opaque to
it. "Where does it explain the Caesar cipher?" only hits if those words are in a
filename.
**The cost:** one `search` walks ~26 course pages, ~30 board skeletons plus card
fetches, and ~150 per-element file listings — roughly **270 upstream requests**,
several seconds, every time.
**Design sketch**
- Postgres `tsvector` + GIN. Use the **`german`** dictionary, so *Verschlüsselung*
matches *Verschlüsselungsverfahren*; the default `english` config stems these
wrongly.
- Add **`pg_trgm`** alongside. German compounds defeat stemming —
*Datenschutzgrundverordnung* will not match *Datenschutz* — and trigram
similarity also absorbs typos.
- Index: titles and text of courses, boards, cards, lessons, tasks; file names;
and **extracted file text**.
- Later, optionally: embeddings in `pgvector`, so "the thing about encryption"
finds *Verschlüsselung* without sharing a word. Needs an embedding model in the
loop, so it is a separate step, not part of this.
## 2. Cache, with bypass — BUILT (`search fresh=true`, `refresh_index`)
**The split that makes staleness tolerable** — index is *discovery*, live API is
*detail*. Search the index to find where something is; always re-fetch it to read
it. A stale index then costs a missed or spurious hit, never wrong content. That
is a far better failure mode than a cache that can serve last week's homework
list as current.
- `fresh: true` on `search` (re-crawl the relevant course first) and on the
detail tools.
- **State freshness in every cached result** ("indexed 2h ago"), regardless of
the flag — a bypass only helps if the agent knows it needs one.
- **Extracted file text is cacheable indefinitely.** A `fileRecord` id maps to
immutable bytes; an edit produces a new record. PDF parsing happens once, ever.
Cheapest large win here, and free of staleness risk.
- Sync: nightly full crawl (~270 requests, trivial) plus on-demand per course.
**Consequence to weigh:** coursework text would then live on the Pi at rest,
outside Schulcloud. It does not weaken the read-only property, but it is new —
consider disk encryption and whether it lands in a backup.
## 3. "What's new since …" — BUILT (`what_changed`)
Falls out of having an index with history: diff successive crawls to surface new
boards, cards, files and tasks. **Impossible today at any speed** — the API has no
changed-since filter anywhere. Arguably the most useful item on this list for a
student, and nearly free once the sync job exists.
## 4. Remaining gaps, measured
Scanned the instance's 98 GET routes against what the tools cover. Almost
everything student-facing is now covered; what is left, with live checks:
| Area | State on this account |
|---|---|
| `GET /groups/class` | **3 real classes** — the only uncovered endpoint with data |
| `GET /rooms` | empty — the feature is unused here |
| `GET /media-boards/me` | exists, no content |
| `GET /course-info` | 403 — teacher/admin only |
| `GET /alert` | empty |
| `tools`, `oauth2`, `school`, `registrations`, `systems`, `user-login-migrations` | admin/auth plumbing, not student-facing |
Genuinely unavailable, not merely uncovered:
- **Collaborative text editor contents** — the endpoint returns an editor URL,
never the document text.
- **Calendar** — a separate service, absent from the v3 document.
- **Image-only PDFs** — 22 of 255 indexed files are scans with no text layer, so
their contents cannot be searched without OCR.
- **Numeric grades** — the API's `grade` was null on every graded submission
here, so that path stays unverified against real data.
## 5. Still open
- **Video/audio transcription** — this account has 5 MP4s and a WebM that are
currently just "here is a file you cannot read".
- **MCP resources and prompts** — expose courses/boards as attachable resources;
canned prompts such as "summarise this week's homework".
- **Calendar** — a separate `schulcloud-calendar` service, not part of the v3 API
mapped in `API.md`.
## 5. Ruled out
- **Collaborative text editor contents.** Confirmed unavailable:
`GET /api/v3/collaborative-text-editor/{parentType}/{parentId}` returns a URL
to the editor, never the document text.
- **OCR — partially wrong, revised.** For *reading*, it remains unnecessary:
images go to Claude inline and it reads them. For *indexing* it is a real gap.
Measured after building the indexer: **3 of 4 sampled course PDFs have no
embedded fonts at all** — they are scans, so extraction legitimately yields
nothing and they are unsearchable by content. `extract.ts` now reports these
as image-only rather than as an empty result. Making them searchable would
need OCR (or page rasterisation plus a vision pass), which is the largest
remaining gap in search coverage.
- **Write tools** (submitting homework, marking tasks done). Technically easy,
but this forfeits the property that makes an internet-facing endpoint
acceptable: that a leaked token cannot act as the user. A deliberate, separate
decision — see the invariant in `CLAUDE.md`.