The endurance test refuted the sliding-window model I committed earlier. A keepalive doing only GET /api/v3/me succeeded at t+0/30/60/90 and was still rejected by t+120 — consistent with the session ending ~2h after LOGIN (t+107), and inconsistent with 2h after the last request, which would have been t+210. This is a live-vs-source divergence, not a misreading: both the current JwtWhitelistAdapter and the legacy Feathers ensureTokenIsWhitelisted re-set the Valkey TTL on every authenticated request, so the source reads as a sliding window. The instance does not behave that way. So the keepalive now calls POST /authentication/refresh-session, the endpoint behind the UI's "Sitzung verlängern" button, which a separate 100s test showed does hold the reported budget at 7200s. It is the only non-GET request in the server: no body, touches only our own session, cannot read or modify user data, and is not exposed as a tool, so no model-driven call can ever be a POST. It logs the returned budget, which makes a failing extension visible before the session is lost. Whether this is sufficient is NOT established. Two mechanisms still fit: an idle TTL that reads fail to refresh (keepalive works), or an absolute cap/revocation anchored at login — e.g. the IDP's back-channel logout, which clears every token for the account rather than one. Added scripts/session-diagnose.mjs to settle it: it logs the budget every 10 min, so a decaying series indicates the former and an abrupt 401 at 7200s the latter. Docs state the open question rather than asserting a mechanism. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
5.6 KiB
Deployment
The shape of it
claude.ai ──HTTPS──▶ VPS (public IP) ──tunnel──▶ Pi 5 (home network)
└─ Caddy ──▶ schulcloud-mcp:8080
│
└──▶ schulcloud-thueringen.de
Claude's custom connectors call the endpoint from Anthropic's cloud, so it must be publicly reachable over real TLS — a localhost tunnel or self-signed cert will not do. The VPS provides the public address; Caddy on the Pi terminates TLS and obtains the certificate.
The container publishes no host port. Caddy reaches it over the shared Docker network, so the only way in from the internet is through Caddy and then through this server's bearer check.
First deploy
git clone <this repo> /opt/schulcloud-mcp
cd /opt/schulcloud-mcp
cp .env.example .env
# Fill in TSC_URL and TSC_JWT_COOKIE (see docs/AUTH.md), then:
openssl rand -hex 32 # → MCP_AUTH_TOKEN
docker compose up -d --build
docker compose logs -f schulcloud-mcp
Expect:
[schulcloud-mcp] listening on 0.0.0.0:8080 — instance https://… , auth enabled, keepalive every 30min
auth DISABLED there means MCP_AUTH_TOKEN is empty — fix it before exposing
the service. A keepalive: token rejected (401) line right after startup means
the Schulcloud token is dead and needs replacing (see docs/AUTH.md); the server
will run but every tool will fail.
Joining the existing Caddy
The Pi already runs Caddy and PostgreSQL in a Compose project. This server needs neither a database nor its own Caddy — only a network it shares with the existing one.
Find the network Caddy is on:
docker inspect -f '{{range $k,$v := .NetworkSettings.Networks}}{{$k}}{{"\n"}}{{end}}' <caddy-container>
Then in docker-compose.yml, set that name and mark it external:
networks:
caddy:
external: true
name: <the network name you just found>
Append deploy/Caddyfile.snippet to the Pi's Caddyfile, replacing
mcp.example.org with the real hostname, and reload:
docker exec <caddy-container> caddy reload --config /etc/caddy/Caddyfile
Two settings in that snippet matter and are easy to miss:
flush_interval -1— MCP's Streamable HTTP transport holds a server-sent-events channel open. Without this, Caddy buffers it and the connector hangs with no error.read_timeout/write_timeoutof 300s — asearchcall walks every course and can take tens of seconds. Caddy's defaults will cut it off.
Ports and DNS
- DNS for the hostname points at the VPS, not the Pi.
- The VPS forwards 80 and 443 to the Pi's Caddy. Port 80 must work too, or Caddy cannot complete the ACME HTTP challenge.
- Nothing else needs to be exposed.
Verifying from outside
curl -s https://mcp.example.org/healthz
# {"status":"ok","sessions":0}
curl -s -o /dev/null -w '%{http_code}\n' -X POST https://mcp.example.org/mcp \
-H 'content-type: application/json' -d '{}'
# 401 ← the bearer check is live
If /healthz answers but /mcp returns 401 with a correct token, check that
the token in .env matches the one in the connector exactly — no trailing
newline from a copy-paste.
Connecting Claude
- claude.ai → Settings → Connectors → Add custom connector.
- URL:
https://mcp.example.org/mcp - Under Advanced settings, add the bearer token as an authorization
header. If your organisation has no header-auth field, the server also
accepts the token as
X-Api-Key. - Enable the connector in a conversation via + → Add connectors.
Ask "which courses am I in?" as a first check — that exercises auth, the Schulcloud token and the API in one call.
Running it locally instead
For Claude Code or Claude Desktop on your own machine, skip all of the above and use stdio:
{
"mcpServers": {
"schulcloud": {
"command": "node",
"args": ["/path/to/schulcloud-mcp/dist/bin/stdio.js"],
"env": {
"TSC_URL": "https://schulcloud-thueringen.de",
"TSC_JWT_COOKIE": "…"
}
}
}
}
MCP_AUTH_TOKEN is irrelevant in stdio mode — there is no network listener.
Updating
cd /opt/schulcloud-mcp && git pull
docker compose up -d --build
docker compose exec schulcloud-mcp node -e "1" # sanity
npm run probe # re-verify the API assumptions
Operational notes
- Restart policy is
unless-stopped; the container comes back after a reboot. - Sessions are in-memory and dropped after 30 minutes idle. A restart invalidates them; Claude re-initializes transparently.
- Logs are capped at 3 × 10 MB. The Authorization header is never logged.
- The container is read-only with
cap_drop: ALLandno-new-privileges, running as the unprivilegednodeuser. It writes nothing to disk — downloads are streamed through memory, capped atMAX_DOWNLOAD_BYTES(25 MiB default). - The Schulcloud session ends ~2 hours after login, and API reads do not
extend it, so the server calls
refresh-sessionevery 30 minutes. Watch forkeepalive: session extended, 7200sin the logs; a falling budget is the early warning. Downtime longer than two hours kills the token and restarting does not recover it — a long power cut means pasting a freshTSC_JWT_COOKIE. Whether the keepalive suffices at all is still being measured; see docs/AUTH.md. - Monthly chore: refresh
TSC_JWT_COOKIEbefore its 30-day hard expiry.npm run probereports both clocks.