You were right that this data never reaches the browser as an API call.
The legacy front end calls the Feathers API server-side for
submission.comment, submission.grade and submission.gradeComment and
renders them into GET /homework/{taskId}. That Feathers API is not
exposed publicly — /api/v1/* 404s — so the rendered page is the only way
to reach these fields from outside.
core/homework-page.ts parses it, hooked on the data-testid attributes
the project's own e2e tests use rather than incidental markup. The page
authenticates by jwt *cookie*; an Authorization header is ignored and
redirects to the identity provider. Every field is optional and parse
failures return undefined, so a markup change degrades to "not found"
and cannot break get_task. The wording distinguishes the two: absent
feedback is reported as not found, never as none given.
Measured on one course: 4 of 7 graded submissions carry feedback no API
call can return — "vollständig und nachvollziehbar", "Feedback siehe
Zettel", and so on.
This exposed a bug in a shared utility: htmlToText decoded only six
entities, so any named entity passed through raw. German content makes
that routine — "vollständig" would have reached the model verbatim
from boards and task descriptions too, not just here. It now decodes
named, decimal and hex references in one pass, so ä stays
literal instead of decoding twice, and leaves unknown names alone rather
than mangling them.
Also fixes a documented-recovery bug found while restoring the session:
`docker compose restart` does not re-read env_file, so it silently kept
serving the dead token. `up -d` is correct and the docs said the wrong
thing.
78 tests, 38/38 smoke; verified end to end through Claude Code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
123 lines
6.1 KiB
Markdown
123 lines
6.1 KiB
Markdown
# Extensions: built and possible
|
||
|
||
Items 1–3 are **now implemented** — see `docs/CLI.md` and the `store/`,
|
||
`indexer/` and `cli/` modules. What remains below is the rationale (worth
|
||
keeping) and the items still open.
|
||
|
||
## Decisions already taken
|
||
|
||
| Decision | Choice | Rejected |
|
||
|---|---|---|
|
||
| Endpoint auth | Static bearer token (`MCP_AUTH_TOKEN`, constant-time check) | OAuth 2.1; secret URL path |
|
||
| State | Stateless — every tool call hits the API live | Postgres cache; cache + full-text search |
|
||
| File delivery | Inline extracted text; images inline | Raw base64 only |
|
||
| **If** persistence is added | **The Pi's existing PostgreSQL, with its own database and user** | A dedicated container; SQLite |
|
||
|
||
Two of the rejected options need no revisiting. OAuth 2.1 is disproportionate
|
||
machinery for a single-user endpoint a bearer token already protects. "Raw bytes
|
||
as well as extracted text" got built anyway — `download_file` takes `raw: true`.
|
||
|
||
## 1. Full-text search over file contents — BUILT
|
||
|
||
**The gap:** `search` covers course/board/card/lesson/task titles and text, and
|
||
file *names* — never file *contents*. The 93 PDFs in this account are opaque to
|
||
it. "Where does it explain the Caesar cipher?" only hits if those words are in a
|
||
filename.
|
||
|
||
**The cost:** one `search` walks ~26 course pages, ~30 board skeletons plus card
|
||
fetches, and ~150 per-element file listings — roughly **270 upstream requests**,
|
||
several seconds, every time.
|
||
|
||
**Design sketch**
|
||
|
||
- Postgres `tsvector` + GIN. Use the **`german`** dictionary, so *Verschlüsselung*
|
||
matches *Verschlüsselungsverfahren*; the default `english` config stems these
|
||
wrongly.
|
||
- Add **`pg_trgm`** alongside. German compounds defeat stemming —
|
||
*Datenschutzgrundverordnung* will not match *Datenschutz* — and trigram
|
||
similarity also absorbs typos.
|
||
- Index: titles and text of courses, boards, cards, lessons, tasks; file names;
|
||
and **extracted file text**.
|
||
- Later, optionally: embeddings in `pgvector`, so "the thing about encryption"
|
||
finds *Verschlüsselung* without sharing a word. Needs an embedding model in the
|
||
loop, so it is a separate step, not part of this.
|
||
|
||
## 2. Cache, with bypass — BUILT (`search fresh=true`, `refresh_index`)
|
||
|
||
**The split that makes staleness tolerable** — index is *discovery*, live API is
|
||
*detail*. Search the index to find where something is; always re-fetch it to read
|
||
it. A stale index then costs a missed or spurious hit, never wrong content. That
|
||
is a far better failure mode than a cache that can serve last week's homework
|
||
list as current.
|
||
|
||
- `fresh: true` on `search` (re-crawl the relevant course first) and on the
|
||
detail tools.
|
||
- **State freshness in every cached result** ("indexed 2h ago"), regardless of
|
||
the flag — a bypass only helps if the agent knows it needs one.
|
||
- **Extracted file text is cacheable indefinitely.** A `fileRecord` id maps to
|
||
immutable bytes; an edit produces a new record. PDF parsing happens once, ever.
|
||
Cheapest large win here, and free of staleness risk.
|
||
- Sync: nightly full crawl (~270 requests, trivial) plus on-demand per course.
|
||
|
||
**Consequence to weigh:** coursework text would then live on the Pi at rest,
|
||
outside Schulcloud. It does not weaken the read-only property, but it is new —
|
||
consider disk encryption and whether it lands in a backup.
|
||
|
||
## 3. "What's new since …" — BUILT (`what_changed`)
|
||
|
||
Falls out of having an index with history: diff successive crawls to surface new
|
||
boards, cards, files and tasks. **Impossible today at any speed** — the API has no
|
||
changed-since filter anywhere. Arguably the most useful item on this list for a
|
||
student, and nearly free once the sync job exists.
|
||
|
||
## 4. Remaining gaps, measured
|
||
|
||
Scanned the instance's 98 GET routes against what the tools cover. Almost
|
||
everything student-facing is now covered; what is left, with live checks:
|
||
|
||
| Area | State on this account |
|
||
|---|---|
|
||
| `GET /groups/class` | **3 real classes** — the only uncovered endpoint with data |
|
||
| `GET /rooms` | empty — the feature is unused here |
|
||
| `GET /media-boards/me` | exists, no content |
|
||
| `GET /course-info` | 403 — teacher/admin only |
|
||
| `GET /alert` | empty |
|
||
| `tools`, `oauth2`, `school`, `registrations`, `systems`, `user-login-migrations` | admin/auth plumbing, not student-facing |
|
||
|
||
Genuinely unavailable, not merely uncovered:
|
||
|
||
- **Collaborative text editor contents** — the endpoint returns an editor URL,
|
||
never the document text.
|
||
- **Calendar** — a separate service, absent from the v3 document.
|
||
- **Image-only PDFs** — 22 of 255 indexed files are scans with no text layer, so
|
||
their contents cannot be searched without OCR.
|
||
- **Numeric grades** — the API's `grade` was null on every graded submission
|
||
here, so that path stays unverified against real data.
|
||
|
||
## 5. Still open
|
||
|
||
- **Video/audio transcription** — this account has 5 MP4s and a WebM that are
|
||
currently just "here is a file you cannot read".
|
||
- **MCP resources and prompts** — expose courses/boards as attachable resources;
|
||
canned prompts such as "summarise this week's homework".
|
||
- **Calendar** — a separate `schulcloud-calendar` service, not part of the v3 API
|
||
mapped in `API.md`.
|
||
|
||
## 5. Ruled out
|
||
|
||
- **Collaborative text editor contents.** Confirmed unavailable:
|
||
`GET /api/v3/collaborative-text-editor/{parentType}/{parentId}` returns a URL
|
||
to the editor, never the document text.
|
||
- **OCR — partially wrong, revised.** For *reading*, it remains unnecessary:
|
||
images go to Claude inline and it reads them. For *indexing* it is a real gap.
|
||
Measured after building the indexer: **3 of 4 sampled course PDFs have no
|
||
embedded fonts at all** — they are scans, so extraction legitimately yields
|
||
nothing and they are unsearchable by content. `extract.ts` now reports these
|
||
as image-only rather than as an empty result. Making them searchable would
|
||
need OCR (or page rasterisation plus a vision pass), which is the largest
|
||
remaining gap in search coverage.
|
||
- **Write tools** (submitting homework, marking tasks done). Technically easy,
|
||
but this forfeits the property that makes an internet-facing endpoint
|
||
acceptable: that a leaked token cannot act as the user. A deliberate, separate
|
||
decision — see the invariant in `CLAUDE.md`.
|