Some checks failed
The decoder died mid-task and it took four separate findings to explain, each of which read as something else: 1. OOM-KILLED, REPORTED AS A CLEAN EXIT. `OOMKilled: true` with **ExitCode 0**. So `--restart on-failure` would treat a memory kill as a successful finish and leave the agent down -- the policy has to be `unless-stopped`. 2. THE JOB CAP WAS SET AND THEN REMOVED THREE LINES LATER. build-reborn has always exported CARGO_BUILD_JOBS, but a raw `cargo test --release -p sylpheed-formats` never reaches the wrapper. Adding `-e CARGO_BUILD_JOBS` to the launcher did not help either: the entrypoint recomputes and exports over it unconditionally. An explicit value now wins, and says so in the log. 3. THE MEMORY CONSTANT WAS WRONG. `mem_gib * 2 / 3` assumes ~1.5 GB per job; release rustc on this workspace needs ~2 GB, and 4 jobs in 6 GB is what died. Divisor is now 2. 4. `--continue` CANNOT RESUME AN ABRUPT DEATH, which is the only kind we get. It resolves through ~/.claude.json's per-project `history`/`lastSessionId`, and MEASURED mid-session both are None -- they are written at a graceful shutdown. A killed container never writes them, so `--continue` answered "No conversation found to continue" with 33 MB of transcripts in the volume beside it. Persisting .claude.json did not help, because the fields were never populated in the first place; that attempt is removed rather than left in looking useful. The TRANSCRIPTS are durable and named by session id, so the entrypoint reads the id off the newest one for its cwd and passes `--resume <id>`. Verified on both agents: each reattached to its exact prior session and appended to the same file rather than opening a new one. The /loop prompt is still passed alongside `--resume`, so the loop is RE-ARMED rather than merely restored -- a resumed conversation with no wake-up scheduled answers once and stops, which looks like resuming and is not. Restarting into the same death is guarded at the other end: a start less than 120 s after the previous one begins FRESH instead of continuing back into whatever killed it. That fired correctly during this work. On resume the agent is told it was restarted, that its in-progress work is uncommitted in the tree, that any build or capture it had running did not finish and its absence is not a result, and which wrapper to prefer over a raw release build.
188 lines
8.1 KiB
Bash
Executable File
188 lines
8.1 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
# Launcher for the Godot port agent.
|
|
#
|
|
# ./sylph-port build build the image
|
|
# ./sylph-port shell interactive shell
|
|
# ./sylph-port loose [task] detached, self-running on a fixed interval
|
|
# ./sylph-port logs -f follow it
|
|
# ./sylph-port attach chat with it (Ctrl-P Ctrl-Q to leave it running)
|
|
# ./sylph-port remote a link to chat with it from anywhere
|
|
# ./sylph-port stop stop it
|
|
#
|
|
# Env:
|
|
# SYLPH_PORT_CPUS / SYLPH_PORT_MEM_GB override the cap (default 3 / 4)
|
|
# SYLPH_PORT_REPO repo to mount at /work (default: this script's parent)
|
|
# SYLPH_DISC extracted disc root
|
|
# SYLPH_GIT_CREDENTIALS file with `https://<user>:<token>@host` for push-work
|
|
# SYLPH_LOOP_INTERVAL fixed loop cadence (default 45m)
|
|
#
|
|
# ── Two hard-won constraints ────────────────────────────────────────────────
|
|
#
|
|
# 1. THIS REPO IS ITS OWN CLONE. It is deliberately NOT the tree the RE agent
|
|
# or a human is working in. Sharing a working tree between two writers means
|
|
# files change under whoever is mid-edit, and a `git add -A` by one sweeps up
|
|
# the other's work. That happened; do not re-create it.
|
|
#
|
|
# 2. IDENTITY GOES IN THE ENVIRONMENT, NOT `.git/config`. Writing `[user]` into
|
|
# a repo's config captures every commit made in that tree, including a
|
|
# human's. GIT_AUTHOR_*/GIT_COMMITTER_* apply to this container's commits and
|
|
# nobody else's.
|
|
set -euo pipefail
|
|
|
|
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
|
# The repo to mount at /work. Overridable so this script can be run from a
|
|
# worktree -- a human editing on `main` must not repoint the agent's checkout.
|
|
# The checkout is now a volume the container clones into, so this is only used
|
|
# to locate things that live BESIDE the repository -- the disc, chiefly. Three
|
|
# levels up from docker/port/ is the workspace root.
|
|
WORKSPACE="$(cd "$HERE/../../.." && pwd)"
|
|
IMAGE="${SYLPH_PORT_IMAGE:-sylpheed-port:latest}"
|
|
NAME="${SYLPH_PORT_NAME:-sylpheed-port}"
|
|
|
|
# Half of what the RE container takes. That container builds a C++ emulator and
|
|
# drives it; this one converts assets and runs Godot. Two full-size containers
|
|
# do not fit on a 12-core / 15 GB box beside a desktop -- memory is the binding
|
|
# constraint, and an over-committed build has crashed this machine before.
|
|
CPUS="${SYLPH_PORT_CPUS:-3}"
|
|
MEM_GB="${SYLPH_PORT_MEM_GB:-4}"
|
|
|
|
DISC="${SYLPH_DISC:-$(cd "$WORKSPACE/sylph_extract" 2>/dev/null && pwd || true)}"
|
|
|
|
docker_args() {
|
|
local _out=(
|
|
--name "$NAME"
|
|
--hostname sylph-port
|
|
--cpus "$CPUS"
|
|
--memory "${MEM_GB}g"
|
|
--memory-swap "${MEM_GB}g" # no swap escape hatch: a swapping build
|
|
# thrashes the whole host
|
|
--pids-limit 2048
|
|
-v "sylpheed-port-repo:/work"
|
|
-v "sylpheed-port-target:/sylph-home/port/target-container"
|
|
# CARGO_HOME on a volume, not the container overlay: without it the pinned
|
|
# decoder source is re-fetched from the network on every fresh container.
|
|
-v "sylpheed-port-cargo:/sylph-home/port/.cargo"
|
|
-v "sylpheed-port-claude:/sylph-home/port/.claude"
|
|
-v "${SYLPH_CLAUDE_HOME:-$HOME/.claude}:/sylph-home/port/.claude.seed:ro"
|
|
-v "${SYLPH_CLAUDE_JSON:-$HOME/.claude.json}:/sylph-home/port/.claude.host.json:ro"
|
|
-v "sylpheed-exchange:/exchange"
|
|
-e "PROJECT_DIR=/work"
|
|
# Same guardrail as the decoder, added the same day and for its reason: the
|
|
# decoder was OOM-killed mid-task by a RAW `cargo test --release`, which
|
|
# never reaches `build-export`/`build-reference-cli` and so never saw their
|
|
# CARGO_BUILD_JOBS. This container is smaller (4 GB, 3 CPUs), so the same
|
|
# bypass is at least as easy to hit here.
|
|
-e "CARGO_BUILD_JOBS=${SYLPH_PORT_JOBS:-2}"
|
|
-e "SYLPH_EXCHANGE=/exchange"
|
|
-e "SYLPH_AGENT=port"
|
|
-e "SYLPH_REPO_URL=https://git.mc02.dev/fabi/Sylpheed.git"
|
|
)
|
|
|
|
|
|
if [ -n "$DISC" ] && [ -d "$DISC" ]; then
|
|
_out+=(-v "$DISC:/disc:ro" -e "SYLPHEED_DISC=/disc")
|
|
else
|
|
echo "==> NOTE: no extracted disc found; the exporter has nothing to read." >&2
|
|
echo " Set SYLPH_DISC to the directory holding dat/ and hidden/." >&2
|
|
fi
|
|
|
|
# Commits are attributed to the port agent, via the environment so that
|
|
# nothing is written into the repository's config. See constraint 2 above.
|
|
_out+=(
|
|
-e "GIT_AUTHOR_NAME=Sylpheed port agent"
|
|
-e "GIT_AUTHOR_EMAIL=port-agent@localhost"
|
|
-e "GIT_COMMITTER_NAME=Sylpheed port agent"
|
|
-e "GIT_COMMITTER_EMAIL=port-agent@localhost"
|
|
)
|
|
|
|
# Mounted as `.host` and copied to a writable file by the entrypoint, exactly
|
|
# like .claude.json. `credential.helper=store` REWRITES its file after a
|
|
# successful auth -- it writes a temp file and renames over the target, and
|
|
# renaming onto a bind-mount point gives EBUSY, which surfaces as
|
|
# `fatal: unable to write credential store: Device or resource busy`.
|
|
#
|
|
# The push still succeeds, which is the actual danger: a `fatal:` line that is
|
|
# routinely wrong teaches the reader to ignore the one that is real. Mounting
|
|
# rw would also silence it, but then the container can clobber the host's
|
|
# credential file; copying cannot.
|
|
local gitcred="${SYLPH_GIT_CREDENTIALS:-$HOME/.sylph-git-credentials}"
|
|
if [ -f "$gitcred" ]; then
|
|
_out+=(-v "$gitcred:/sylph-home/port/.git-credentials.host:ro")
|
|
else
|
|
echo "==> NOTE: no git credentials at $gitcred — the agent cannot push," >&2
|
|
echo " so its work dies with the container." >&2
|
|
fi
|
|
|
|
[ -n "${ANTHROPIC_API_KEY:-}" ] && _out+=(-e "ANTHROPIC_API_KEY=$ANTHROPIC_API_KEY")
|
|
printf '%s\n' "${_out[@]}"
|
|
}
|
|
|
|
mapfile -t ARGS < <(docker_args)
|
|
|
|
case "${1:-}" in
|
|
build)
|
|
exec docker build -t "$IMAGE" \
|
|
--build-arg "AGENT_UID=$(id -u)" --build-arg "AGENT_GID=$(id -g)" "$HERE"
|
|
;;
|
|
|
|
shell)
|
|
TTY=(-i); [ -t 0 ] && TTY=(-it)
|
|
exec docker run --rm "${TTY[@]}" "${ARGS[@]}" "$IMAGE" bash
|
|
;;
|
|
|
|
loose)
|
|
shift
|
|
TASK="${1:-}"
|
|
if [ -z "$TASK" ]; then
|
|
if [ -f "$HERE/../../docs/agents/port-loop.md" ]; then
|
|
TASK="$(cat "$HERE/../../docs/agents/port-loop.md")"
|
|
else
|
|
TASK="Work the milestones in docs/MISSION.md."
|
|
fi
|
|
fi
|
|
# A FIXED interval, not self-pacing: the one thing an agent deep in a
|
|
# milestone reliably forgets is the bookkeeping after it, and a forgotten
|
|
# wake-up silently ends the loop.
|
|
INTERVAL="${SYLPH_LOOP_INTERVAL-45m}"
|
|
echo "==> loose | cpus=$CPUS mem=${MEM_GB}g pacing=${INTERVAL:-self}"
|
|
echo "==> repo: own clone in volume sylpheed-port-repo -> /work"
|
|
# `unless-stopped`, NOT `on-failure`: an OOM kill on this setup reports
|
|
# `OOMKilled: true` with **ExitCode 0**, so `on-failure` would read a memory
|
|
# kill as a clean finish and leave the agent down. Restarting into the same
|
|
# death is handled in the entrypoint, which refuses to `--continue` when the
|
|
# last start was under two minutes ago.
|
|
docker run -d -i -t --restart unless-stopped "${ARGS[@]}" -e SYLPH_AUTONOMOUS=1 -w /work "$IMAGE" \
|
|
"/loop ${INTERVAL:+$INTERVAL }$TASK" >/dev/null
|
|
echo
|
|
echo " running detached as '$NAME'."
|
|
echo " ./sylph-port remote link to chat with it from anywhere"
|
|
echo " ./sylph-port logs -f follow it"
|
|
echo " ./sylph-port attach chat with it locally"
|
|
echo " ./sylph-port stop stop it"
|
|
;;
|
|
|
|
logs) shift; exec docker logs "$@" "$NAME" ;;
|
|
attach) exec docker attach "$NAME" ;;
|
|
stop) exec docker rm -f "$NAME" ;;
|
|
|
|
remote)
|
|
echo "waiting for the session to register" >&2
|
|
for _ in $(seq 60); do
|
|
# Read the container LOG, not the session transcript. The transcript
|
|
# records every command run inside the container -- including this
|
|
# lookup -- so grepping it matched our own pattern string back.
|
|
url=$(docker logs "$NAME" 2>&1 \
|
|
| grep -aoE 'https://claude\.ai/code/session_[A-Za-z0-9]+' \
|
|
| tail -1 || true)
|
|
[ -n "$url" ] && { echo "$url"; exit 0; }
|
|
sleep 2
|
|
done
|
|
echo "no session link yet — try ./sylph-port logs -f" >&2
|
|
exit 1
|
|
;;
|
|
|
|
*)
|
|
sed -n '2,25p' "$0" | sed 's/^# \{0,1\}//'
|
|
;;
|
|
esac
|