Files
Schulcloud-MCP/docs/DEPLOYMENT.md
MechaCat02 d657ece436 Fix session lifetime: 2h sliding idle timeout, not 30 days
The JWT's exp claim says 30 days, and I took that as the session
lifetime. It is only an outer ceiling. The server also keeps a per-token
whitelist entry in Valkey (jwt:{accountId}:{jti}) whose TTL is
JWT_TIMEOUT_SECONDS — 7200s on this instance — and JwtStrategy.validate
re-sets it on every authenticated request. Two hours idle and the token
is rejected with 29 days still on exp.

Proven, not inferred: the token from yesterday returned 401 at 13.8h old.
The live instance publishes the values unauthenticated at
GET /api/v3/config/public — JWT_TIMEOUT_SECONDS 7200,
JWT_SHOW_TIMEOUT_WARNING_SECONDS 3600, the latter being exactly the
one-hour UI prompt that prompted this investigation.

refresh-session turns out not to be special: it extends through the same
guard as any other route, and uniquely only in returning the remaining
TTL. So the keepalive uses GET /api/v3/me instead, and the server stays
GET-only; the one POST in the repo is in scripts/probe.mjs, where it
reports the idle budget.

JWT_EXTENDED_TIMEOUT_SECONDS (~1 month) exists in the config schema but
is vestigial: privateDevice has no references in the current NestJS
source, and generateJwtAndAddToWhitelist never overrides the TTL.

Also fixes a real breakage this surfaced: TypeScript parameter
properties are rejected by Node's type stripping, so `npm run dev` and
`npm test` both failed on any file reaching them. Rewritten as explicit
fields, and noted in CLAUDE.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 13:07:22 +02:00

5.5 KiB
Raw Blame History

Deployment

The shape of it

claude.ai  ──HTTPS──▶  VPS (public IP)  ──tunnel──▶  Pi 5 (home network)
                                                      └─ Caddy ──▶ schulcloud-mcp:8080
                                                                      │
                                                                      └──▶ schulcloud-thueringen.de

Claude's custom connectors call the endpoint from Anthropic's cloud, so it must be publicly reachable over real TLS — a localhost tunnel or self-signed cert will not do. The VPS provides the public address; Caddy on the Pi terminates TLS and obtains the certificate.

The container publishes no host port. Caddy reaches it over the shared Docker network, so the only way in from the internet is through Caddy and then through this server's bearer check.

First deploy

git clone <this repo> /opt/schulcloud-mcp
cd /opt/schulcloud-mcp

cp .env.example .env
# Fill in TSC_URL and TSC_JWT_COOKIE (see docs/AUTH.md), then:
openssl rand -hex 32   # → MCP_AUTH_TOKEN

docker compose up -d --build
docker compose logs -f schulcloud-mcp

Expect:

[schulcloud-mcp] listening on 0.0.0.0:8080 — instance https://… , auth enabled, keepalive every 30min

auth DISABLED there means MCP_AUTH_TOKEN is empty — fix it before exposing the service. A keepalive: token rejected (401) line right after startup means the Schulcloud token is dead and needs replacing (see docs/AUTH.md); the server will run but every tool will fail.

Joining the existing Caddy

The Pi already runs Caddy and PostgreSQL in a Compose project. This server needs neither a database nor its own Caddy — only a network it shares with the existing one.

Find the network Caddy is on:

docker inspect -f '{{range $k,$v := .NetworkSettings.Networks}}{{$k}}{{"\n"}}{{end}}' <caddy-container>

Then in docker-compose.yml, set that name and mark it external:

networks:
  caddy:
    external: true
    name: <the network name you just found>

Append deploy/Caddyfile.snippet to the Pi's Caddyfile, replacing mcp.example.org with the real hostname, and reload:

docker exec <caddy-container> caddy reload --config /etc/caddy/Caddyfile

Two settings in that snippet matter and are easy to miss:

  • flush_interval -1 — MCP's Streamable HTTP transport holds a server-sent-events channel open. Without this, Caddy buffers it and the connector hangs with no error.
  • read_timeout/write_timeout of 300s — a search call walks every course and can take tens of seconds. Caddy's defaults will cut it off.

Ports and DNS

  • DNS for the hostname points at the VPS, not the Pi.
  • The VPS forwards 80 and 443 to the Pi's Caddy. Port 80 must work too, or Caddy cannot complete the ACME HTTP challenge.
  • Nothing else needs to be exposed.

Verifying from outside

curl -s https://mcp.example.org/healthz
# {"status":"ok","sessions":0}

curl -s -o /dev/null -w '%{http_code}\n' -X POST https://mcp.example.org/mcp \
  -H 'content-type: application/json' -d '{}'
# 401   ← the bearer check is live

If /healthz answers but /mcp returns 401 with a correct token, check that the token in .env matches the one in the connector exactly — no trailing newline from a copy-paste.

Connecting Claude

  1. claude.ai → Settings → Connectors → Add custom connector.
  2. URL: https://mcp.example.org/mcp
  3. Under Advanced settings, add the bearer token as an authorization header. If your organisation has no header-auth field, the server also accepts the token as X-Api-Key.
  4. Enable the connector in a conversation via + → Add connectors.

Ask "which courses am I in?" as a first check — that exercises auth, the Schulcloud token and the API in one call.

Running it locally instead

For Claude Code or Claude Desktop on your own machine, skip all of the above and use stdio:

{
  "mcpServers": {
    "schulcloud": {
      "command": "node",
      "args": ["/path/to/schulcloud-mcp/dist/bin/stdio.js"],
      "env": {
        "TSC_URL": "https://schulcloud-thueringen.de",
        "TSC_JWT_COOKIE": "…"
      }
    }
  }
}

MCP_AUTH_TOKEN is irrelevant in stdio mode — there is no network listener.

Updating

cd /opt/schulcloud-mcp && git pull
docker compose up -d --build
docker compose exec schulcloud-mcp node -e "1" # sanity
npm run probe                                  # re-verify the API assumptions

Operational notes

  • Restart policy is unless-stopped; the container comes back after a reboot.
  • Sessions are in-memory and dropped after 30 minutes idle. A restart invalidates them; Claude re-initializes transparently.
  • Logs are capped at 3 × 10 MB. The Authorization header is never logged.
  • The container is read-only with cap_drop: ALL and no-new-privileges, running as the unprivileged node user. It writes nothing to disk — downloads are streamed through memory, capped at MAX_DOWNLOAD_BYTES (25 MiB default).
  • The Schulcloud session dies after 2 hours of inactivity, so the server pings /api/v3/me every 30 minutes to hold it open. This has a consequence worth planning for: downtime longer than two hours kills the token, and restarting does not recover it — a long power cut means pasting a fresh TSC_JWT_COOKIE. The startup log says so immediately. Background in docs/AUTH.md.
  • Monthly chore: refresh TSC_JWT_COOKIE before its 30-day hard expiry. npm run probe reports both clocks.