The last structural fix for the collision class that has bitten three times. Both containers now clone the repository into their OWN named volume instead of bind-mounting a human's working tree, so an agent's local git config cannot capture a human's commits, a credential helper cannot leak a container-only path onto the host, and a `git add -A` cannot sweep another party's in-flight files. Cloned once at startup and never auto-pulled: pulling under a running agent moves files out from under whatever it is mid-edit, which is the same bug again. Accepted knowingly: Claude Code keys per-project memory off the working directory, so moving off the host path starts that memory empty. The corpus in docs/ is the memory that matters and it travels with the clone. Other changes: * docker/agent -> docker/decoder; the launcher is sylph-decoder. Roles, not "the agent", now that there is more than one. * /reborn is gone -- one repository now, so the port reads HANDOFF from its own checkout rather than through a live read-only mount of someone else's tree. * Canary mounts separately at /canary; it stays a fork tracking upstream. * A shared `sylpheed-exchange` volume at /exchange, with tools/ on PATH so `share` is available in both. * The decoder's credential file gets the .host-copy treatment the port already had -- `credential.helper=store` rewrites by rename-over-target, which is EBUSY on a bind mount and reports a fatal that is not one. * Budget split deliberately: decoder 5 cpu / 6 GB, port 3 / 4, leaving room for the planned Referee. "Half the host" was right when there was one agent. Prompts move to docs/agents/ and are rewritten around the protocol: the oracle is the running game, dynamic RE stays with the decoder, each iteration must attempt to refute one claim of the other, and neither may verify its way out of its own role.
252 lines
12 KiB
Markdown
252 lines
12 KiB
Markdown
# The RE agent container
|
||
|
||
A container an autonomous Claude Code agent can be turned loose in: it builds
|
||
and runs **both** halves of the project — Xenia Canary as the behaviour oracle
|
||
and Sylpheed Reborn as the port — and carries the dynamic-RE toolkit that drives
|
||
the emulator, reads its guest memory and photographs its screen.
|
||
|
||
```bash
|
||
./sylph-agent build # build the image
|
||
./sylph-agent doctor # prove it can do the four things it exists for
|
||
./sylph-agent shell # poke around
|
||
./sylph-agent agent # Claude Code, --dangerously-skip-permissions
|
||
./sylph-agent loose # turn it loose: detached, /loop, self-paced
|
||
./sylph-agent remote # link to chat with it from anywhere
|
||
./sylph-agent attach # chat with it locally
|
||
./sylph-agent logs -f # watch it
|
||
./sylph-agent stop # stop it
|
||
```
|
||
|
||
## On the loose
|
||
|
||
`./sylph-agent loose` starts Claude Code **detached**, with
|
||
`--dangerously-skip-permissions`, running `/loop` on the task in
|
||
[`loop-task.md`](loop-task.md) — work the RE backlog one item at a time, commit
|
||
to `auto/*` branches, never push, record withdrawn results rather than deleting
|
||
them. Pass your own task as an argument, or set `SYLPH_LOOP_INTERVAL=30m` for a
|
||
fixed cadence instead of letting it self-pace.
|
||
|
||
It runs `-d` **without** `--rm`, so the transcript survives the container
|
||
exiting — for an unattended run that is the only record of what happened.
|
||
|
||
### Talking to it
|
||
|
||
Two channels, both live while it works:
|
||
|
||
* **`./sylph-agent remote`** prints a `https://claude.ai/code/session_…` link.
|
||
`loose` starts Claude Code with `--remote-control`, so the session registers
|
||
with your account and you can chat with it from claude.ai or your phone —
|
||
which is the point of a detached run. Disable with `SYLPH_REMOTE=0`, rename
|
||
with `SYLPH_REMOTE_NAME`.
|
||
|
||
Registration takes a minute or two after launch, so `remote` waits for it
|
||
rather than reporting "not found" to what is really "not yet".
|
||
|
||
* **`./sylph-agent attach`** joins the container's own terminal. Type to talk to
|
||
it; **Ctrl-P Ctrl-Q** detaches and leaves it running. Do not press Ctrl-C —
|
||
that goes to the agent.
|
||
|
||
The pty is forced to 200×50 (`stty_init` in `bin/claude-autonomous`). A detached
|
||
`docker run -t` is 80×24, and Claude Code hard-wraps to the terminal width,
|
||
which truncated the Remote Control URL to `…/session_01…` in the one place you
|
||
need to read it — and made `docker logs` almost unreadable besides.
|
||
|
||
**It cannot push.** No git credentials are mounted, deliberately: a human
|
||
reviews before anything leaves the box. Review with
|
||
`git -C <project>/Syplheed-Reborn log --oneline main..auto/<topic>`.
|
||
|
||
### The four gates
|
||
|
||
Claude Code has four one-time prompts, and each one is a silent, permanent hang
|
||
for an agent with nobody at the keyboard — no error, no log line, just a
|
||
container that looks healthy and does nothing. All four are handled:
|
||
|
||
| gate | how |
|
||
|---|---|
|
||
| theme picker | `hasCompletedOnboarding` + `lastOnboardingVersion` in `~/.claude.json` |
|
||
| "do you trust this folder?" | `projects.<path>.hasTrustDialogAccepted` |
|
||
| Bypass Permissions disclaimer | answered in a pty by [`bin/claude-autonomous`](bin/claude-autonomous) — it has no config key, by design |
|
||
| fullscreen-renderer upsell | `fullscreenUpsellSeenCount`, because it fires *mid-session*, after the pty wrapper has handed over |
|
||
|
||
Config keys were read out of the shipped binary's own strings rather than
|
||
guessed. The pty wrapper matches **single words**: Claude Code draws its UI with
|
||
absolute-column escapes between words, so `Yes, I accept` arrives as
|
||
`Yes,\x1b[13GI\x1b[15Gaccept` and a multi-word pattern never matches — which
|
||
looks exactly like the wrapper not running at all. It stops matching once the
|
||
session is live, so nothing later can be answered by accident.
|
||
|
||
## The resource cap
|
||
|
||
The container gets **half the machine**, computed at launch so it stays half on
|
||
any box:
|
||
|
||
| | how |
|
||
|---|---|
|
||
| CPU | `--cpus $(nproc)/2` |
|
||
| memory | `--memory` = half `MemTotal`, **`--memory-swap` equal to it** |
|
||
| `/dev/shm` | a third of the memory cap, min 1 GiB, mounted **exec** |
|
||
| build jobs | derived *inside* the container from **available memory**, not cores |
|
||
|
||
Two of those deserve a word.
|
||
|
||
**No swap headroom.** `--memory-swap` is set equal to `--memory`, so the
|
||
container cannot swap. That is deliberate: a swapping build thrashes the whole
|
||
host, which is precisely the failure the cap exists to prevent. A build that
|
||
would have swapped gets OOM-killed inside the container instead, and the host
|
||
stays usable.
|
||
|
||
**`/dev/shm` is not incidental.** Xenia backs the guest address space with
|
||
`/dev/shm/xenia_memory_*`, and the whole live-memory toolkit (`gmem.py`,
|
||
`gpoke.py`, `mission_state.py`) reads it from there. Docker's default is 64 MiB,
|
||
which is far too small for a 512 MiB console — and it fails as an obscure mmap
|
||
error rather than an out-of-space message.
|
||
|
||
**Build parallelism is memory-bound.** A full-parallel build of this tree has
|
||
OOM-killed the host outright, so the entrypoint computes jobs from *available
|
||
memory* (≈1.5 GiB per C++ TU) and exports it as `SYLPH_JOBS`, `CARGO_BUILD_JOBS`
|
||
and `CMAKE_BUILD_PARALLEL_LEVEL`. Override with `SYLPH_CPUS` / `SYLPH_MEM_GB`.
|
||
|
||
## Inside
|
||
|
||
| command | what |
|
||
|---|---|
|
||
| `build-canary [Release\|Debug]` | configure + build Canary |
|
||
| `build-reborn [build\|test\|ci]` | build/test Reborn, with the disc env wired up |
|
||
| `run-canary [flags…]` | launch Canary with the settings this title needs |
|
||
| `screenshot [out.png]` | grab the display |
|
||
| `sylph-doctor` | self-check |
|
||
| `tools/re-capture/*` | the RE toolkit, already on `PATH` |
|
||
|
||
Layout: project at `/work`, `HOME=/sylph-home/re`, `DISPLAY=:98` — the values
|
||
`tools/re-capture/*.sh` already assume, so the existing toolkit runs unmodified.
|
||
|
||
Build outputs live **outside** the bind mount (`CARGO_TARGET_DIR`,
|
||
`XENIA_BUILD_DIR`, both named Docker volumes). The host builds the same trees,
|
||
and sharing `target/` or `build/` makes host and container reconfigure and
|
||
relink everything the other just did.
|
||
|
||
## Screenshots
|
||
|
||
Two layers, and the distinction matters:
|
||
|
||
* `/usr/local/bin/screenshot` — raw full-root PNG (ImageMagick, falling back to
|
||
ffmpeg's x11grab, then xwd).
|
||
* `tools/re-capture/bin/screenshot` — **first on `PATH`**, wraps the above and
|
||
crops to the *game surface*.
|
||
|
||
The crop is not cosmetic. Xenia's window is a GTK window whose menu bar pushes
|
||
the 1280×720 game image down ~25 px, and every pixel oracle in the toolkit was
|
||
measured against the bare game image. When that offset was unaccounted for, one
|
||
run sat 300 s in front of a plainly visible MAIN MENU reporting "no main menu".
|
||
The wrapper derives the offset from the window's own height rather than a
|
||
per-display constant.
|
||
|
||
For finding a screen at all, prefer `screen_id.py`, which classifies by
|
||
whole-image statistics instead of named pixels.
|
||
|
||
## Vulkan
|
||
|
||
`mesa-vulkan-drivers` + `vulkan-tools` are installed, so Vulkan works with **no
|
||
host GPU** via lavapipe (software — correct, slow). When the host has
|
||
`/dev/dri`, the launcher passes the device through and adds the host's `render`
|
||
and `video` GIDs, and the entrypoint uses the hardware ICD. Force software with
|
||
`SYLPH_VULKAN=sw`. `vulkaninfo --summary` (or `sylph-doctor`) says which you got — and the
|
||
entrypoint reports the device that **actually enumerated**, not the one it asked
|
||
for, because "I passed `/dev/dri`" and "I have hardware Vulkan" are different
|
||
claims.
|
||
|
||
⚠️ **On an NVIDIA host, `/dev/dri` alone does nothing** — Mesa cannot drive an
|
||
NVIDIA card and the proprietary userspace lives outside the image. You need the
|
||
NVIDIA Container Toolkit; the launcher detects the situation and tells you the
|
||
three commands. Until then Canary runs on lavapipe, which is correct but has not
|
||
been observed to reach a rendered frame in a couple of minutes — everything
|
||
*else* (guest memory, the JIT, the live-memory toolkit) works fine on it.
|
||
|
||
## Input, and why there is no virtual gamepad
|
||
|
||
`run-canary` passes `--hid=file --pad_file=/tmp/xenia_pad.txt`; drive it with
|
||
`tools/re-capture/pad.py`. There is deliberately **no `/dev/uinput`**: input
|
||
devices are not namespaced, so a virtual pad created in a container registers
|
||
with the *host's* input stack and every scripted press leaks onto the user's
|
||
desktop.
|
||
|
||
The trap that wasted a session: 360 menus poll `XamInputGetKeystrokeEx`, not
|
||
`GetState` — with `GetKeystroke` stubbed the pad looks completely dead on a
|
||
title screen while its own log shows the press arriving.
|
||
|
||
## Settings that are requirements, not preferences
|
||
|
||
`run-canary` bakes these in; changing them will cost you an afternoon.
|
||
|
||
* **`--apu=sdl` with `SDL_AUDIODRIVER=dummy`** — and **no `--audio` flag**,
|
||
which is not a cvar here (see below). There is no PulseAudio, so `--apu=nop`
|
||
looks like the safe muted choice. It is not: the log fills with
|
||
`CreateDriver failed for index=0`, the guest never gets past the intro movie,
|
||
and the window stays black for 8+ minutes. SDL against a dummy device is
|
||
silent *and* lets the title advance.
|
||
* **One emulator at a time**, enforced with a lockfile. Two at once perturbs
|
||
both and the box.
|
||
* Stale `/dev/shm/xenia_memory_*` from a killed run is removed at launch —
|
||
otherwise the memory readers find two candidates and pick the dead one.
|
||
|
||
## Claude Code
|
||
|
||
Runs as an unprivileged `agent` user, because `--dangerously-skip-permissions`
|
||
is **refused under root**. `./sylph-agent agent` sets `SYLPH_AUTONOMOUS=1` and
|
||
the entrypoint adds the flag.
|
||
|
||
Auth comes from the host `~/.claude`, bind-mounted read-write (token refresh
|
||
needs to write). **That directory also holds your memory and project state**, so
|
||
the container agent and you share it. Point `SYLPH_CLAUDE_HOME` at a separate
|
||
directory to isolate it, or set `ANTHROPIC_API_KEY` instead.
|
||
|
||
`~/.claude.json` is different: mounted **read-only** at a staging path and
|
||
copied in, so the container cannot rewrite your host config — and so a version
|
||
skew between the container's Claude Code and yours cannot re-trigger onboarding.
|
||
|
||
The project is bind-mounted **twice**, at `/work` and at its own host path. The
|
||
host path is what makes memory carry over: Claude Code derives its per-project
|
||
state key from the working directory, so running at `/work` would hand the agent
|
||
an empty project instead of the accumulated one. Verified — a loose run reports
|
||
`MEMORY=yes` and reads back the same branch and backlog you see.
|
||
|
||
## Host prerequisites
|
||
|
||
* **A Vulkan SDK** (LunarG), for *building* only. Canary's shader step calls
|
||
`spirv-opt --canonicalize-ids`, which Ubuntu's packaged SPIRV-Tools (v2025.1)
|
||
does not have — the build then dies ~500 objects in, and the error you see is
|
||
a Python `TypeError`, not the real message. The launcher mounts the host's SDK
|
||
read-only at its own path and sets `VULKAN_SDK`; that also guarantees the
|
||
container produces byte-identical shaders to a host build.
|
||
* **`nvidia-container-toolkit`**, for hardware Vulkan — see below.
|
||
|
||
## Things that will waste your afternoon
|
||
|
||
Each of these was hit while bringing this container up.
|
||
|
||
* **An unknown xenia flag hangs; it does not error.** `ParseLaunchArguments`
|
||
calls `ShowSimpleMessageBox` *before logging is initialised*, and that SDL
|
||
dialog blocks on `XIfEvent` forever with nobody to click it. The symptom is a
|
||
10×10 window, a completely empty log and no guest memory — which reads like a
|
||
hang deep in the emulator. `--audio` is **not** a cvar in this tree despite
|
||
appearing in the RE notes; `--apu=sdl` is the real one. If Canary appears to
|
||
hang at startup, suspect a typo'd flag first.
|
||
* **`/dev/shm` must be `exec`.** Docker mounts it `noexec`, and xenia maps its
|
||
JIT code cache out of a shm file. With `noexec` it dies at startup with
|
||
"Unable to allocate code cache generated code storage / Cannot initalize
|
||
processor", which reads like an address-space clash. The launcher uses
|
||
`--tmpfs /dev/shm:rw,exec,…` rather than `--shm-size`.
|
||
* **gdb needs root inside the container.** `--cap-add SYS_PTRACE` is passed, but
|
||
the *host's* `kernel.yama.ptrace_scope=1` still blocks attaching to a
|
||
non-descendant. Use `sudo gdb -p <pid>` (passwordless), or launch the target
|
||
under gdb so it is a child.
|
||
* **Named volumes need their mount points to exist in the image**, or Docker
|
||
creates them root-owned and the first write fails obscurely.
|
||
|
||
## Known limitation
|
||
|
||
`build-reborn ci` runs the native legs only. `just ci`'s wasm check does not
|
||
build, for a pre-existing reason unrelated to any change under test: the
|
||
workspace pins `tokio = { features = ["full"] }`, which pulls `mio`, which
|
||
refuses to compile for `wasm32-unknown-unknown`.
|