Some checks failed
The decoder died mid-task and it took four separate findings to explain, each of which read as something else: 1. OOM-KILLED, REPORTED AS A CLEAN EXIT. `OOMKilled: true` with **ExitCode 0**. So `--restart on-failure` would treat a memory kill as a successful finish and leave the agent down -- the policy has to be `unless-stopped`. 2. THE JOB CAP WAS SET AND THEN REMOVED THREE LINES LATER. build-reborn has always exported CARGO_BUILD_JOBS, but a raw `cargo test --release -p sylpheed-formats` never reaches the wrapper. Adding `-e CARGO_BUILD_JOBS` to the launcher did not help either: the entrypoint recomputes and exports over it unconditionally. An explicit value now wins, and says so in the log. 3. THE MEMORY CONSTANT WAS WRONG. `mem_gib * 2 / 3` assumes ~1.5 GB per job; release rustc on this workspace needs ~2 GB, and 4 jobs in 6 GB is what died. Divisor is now 2. 4. `--continue` CANNOT RESUME AN ABRUPT DEATH, which is the only kind we get. It resolves through ~/.claude.json's per-project `history`/`lastSessionId`, and MEASURED mid-session both are None -- they are written at a graceful shutdown. A killed container never writes them, so `--continue` answered "No conversation found to continue" with 33 MB of transcripts in the volume beside it. Persisting .claude.json did not help, because the fields were never populated in the first place; that attempt is removed rather than left in looking useful. The TRANSCRIPTS are durable and named by session id, so the entrypoint reads the id off the newest one for its cwd and passes `--resume <id>`. Verified on both agents: each reattached to its exact prior session and appended to the same file rather than opening a new one. The /loop prompt is still passed alongside `--resume`, so the loop is RE-ARMED rather than merely restored -- a resumed conversation with no wake-up scheduled answers once and stops, which looks like resuming and is not. Restarting into the same death is guarded at the other end: a start less than 120 s after the previous one begins FRESH instead of continuing back into whatever killed it. That fired correctly during this work. On resume the agent is told it was restarted, that its in-progress work is uncommitted in the tree, that any build or capture it had running did not finish and its absence is not a result, and which wrapper to prefer over a raw release build.
197 lines
10 KiB
Bash
Executable File
197 lines
10 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
# Bring up the headless display, then hand over.
|
|
#
|
|
# Xvfb and openbox are started as children of PID 1 (tini), NOT of the agent's
|
|
# shell, so they outlive any single command. The RE container learned this the
|
|
# hard way: a display owned by a shell gets reaped when that shell exits, which
|
|
# reads as "Xvfb dies on its own every few minutes".
|
|
set -euo pipefail
|
|
|
|
: "${DISPLAY:=:97}"
|
|
: "${SCREEN_GEOMETRY:=1280x720x24}"
|
|
|
|
if ! xdpyinfo -display "$DISPLAY" >/dev/null 2>&1; then
|
|
Xvfb "$DISPLAY" -screen 0 "$SCREEN_GEOMETRY" -nolisten tcp &
|
|
for _ in $(seq 50); do
|
|
xdpyinfo -display "$DISPLAY" >/dev/null 2>&1 && break
|
|
sleep 0.1
|
|
done
|
|
openbox >/dev/null 2>&1 &
|
|
fi
|
|
echo "[entrypoint] display $DISPLAY ready ($SCREEN_GEOMETRY)"
|
|
|
|
|
|
# ── Claude state: this agent's own, seeded once from the host ───────────────
|
|
# Isolated per agent. Both working directories are /work, and Claude Code keys
|
|
# its per-project state off the working directory -- so a SHARED ~/.claude put
|
|
# two independent agents in the same projects/-work/ directory, which undoes the
|
|
# point of giving them separate checkouts.
|
|
#
|
|
# Seeded rather than shared because credentials live in .credentials.json and a
|
|
# token refresh must be able to write. Copying once means each agent refreshes
|
|
# its own token and neither can corrupt the host's.
|
|
# Re-seed whenever the HOST's credentials are newer than ours, not only when
|
|
# ours are missing. The missing-only guard meant an expired token could never be
|
|
# replaced: the file existed, so the copy was skipped, and restarting the
|
|
# container changed nothing. A human re-logging in on the host is exactly the
|
|
# recovery path, and it has to reach here.
|
|
#
|
|
# Newer-wins rather than always-copy, because the container refreshes its own
|
|
# token during a run and that copy may legitimately be the fresher one.
|
|
if [ -d "$HOME/.claude.seed" ] && \
|
|
{ [ ! -s "$HOME/.claude/.credentials.json" ] || \
|
|
[ "$HOME/.claude.seed/.credentials.json" -nt "$HOME/.claude/.credentials.json" ]; }; then
|
|
mkdir -p "$HOME/.claude"
|
|
cp -a "$HOME/.claude.seed/.credentials.json" "$HOME/.claude/" 2>/dev/null || true
|
|
for f in settings.json CLAUDE.md; do
|
|
[ -e "$HOME/.claude.seed/$f" ] && cp -a "$HOME/.claude.seed/$f" "$HOME/.claude/" 2>/dev/null || true
|
|
done
|
|
echo "[entrypoint] refreshed ~/.claude credentials from the host"
|
|
fi
|
|
|
|
# Seed ~/.claude.json from the host's read-only copy, then stamp onboarding as
|
|
# complete. Claude Code re-runs its first-run wizard whenever
|
|
# lastOnboardingVersion differs from the installed version, so a container with a
|
|
# newer Claude than the host stops on the theme picker -- no error, no log line,
|
|
# and an unattended agent sits there forever.
|
|
if [ -f "$HOME/.claude.host.json" ] && [ ! -s "$HOME/.claude.json" ]; then
|
|
cp "$HOME/.claude.host.json" "$HOME/.claude.json" 2>/dev/null || true
|
|
fi
|
|
|
|
# Nothing is restored into `~/.claude.json` on purpose. Resuming is done by
|
|
# SESSION ID off the transcript instead — see the resume block below for why the
|
|
# index cannot serve.
|
|
# Same reason as .claude.json above: `credential.helper=store` rewrites this
|
|
# file by rename-over-target, which fails with EBUSY on a bind mount. Copy it to
|
|
# a writable path; nothing is ever written back to the host's file.
|
|
if [ -f "$HOME/.git-credentials.host" ]; then
|
|
cp "$HOME/.git-credentials.host" "$HOME/.git-credentials" 2>/dev/null || true
|
|
chmod 600 "$HOME/.git-credentials" 2>/dev/null || true
|
|
fi
|
|
|
|
CLAUDE_VER=$(claude --version 2>/dev/null | grep -oE '^[0-9][0-9.]*' || echo 0.0.0)
|
|
python3 /usr/local/bin/seed-claude-config.py "$HOME/.claude.json" "$CLAUDE_VER" \
|
|
"$PWD" "${PROJECT_DIR:-/work}" "$HOME" || true
|
|
chmod 600 "$HOME/.claude.json" 2>/dev/null || true
|
|
|
|
# ── The repository, cloned into THIS AGENT'S OWN volume ─────────────────────
|
|
# Not a bind mount of a human's working tree. That arrangement bit this project
|
|
# three times: an agent's `git config --local` captured a human's commits, a
|
|
# credential helper leaked a container-only path onto the host, and a `git add
|
|
# -A` swept an agent's in-flight files into somebody else's commit. Separate
|
|
# checkouts make all three impossible rather than discouraged.
|
|
#
|
|
# Cloned ONCE. Never auto-pulled: pulling under a running agent moves files out
|
|
# from under whatever it is mid-edit, which is the same class of bug again.
|
|
if ! git -C /work rev-parse --verify HEAD >/dev/null 2>&1; then
|
|
# Checks for a usable HEAD, not merely a .git directory. A clone interrupted
|
|
# partway -- the container stopped while it ran, which has happened -- leaves
|
|
# a .git with no commits, and a presence check would then skip the retry
|
|
# forever and hand the agent an empty repository.
|
|
echo "[entrypoint] cloning ${SYLPH_REPO_URL:-https://git.mc02.dev/fabi/Sylpheed.git} into /work"
|
|
_tmp=$(mktemp -d)
|
|
if git clone --quiet "${SYLPH_REPO_URL:-https://git.mc02.dev/fabi/Sylpheed.git}" "$_tmp/r"; then
|
|
find /work -mindepth 1 -maxdepth 1 -exec rm -rf {} + 2>/dev/null || true
|
|
mv "$_tmp/r"/.[!.]* "$_tmp/r"/* /work/ 2>/dev/null || true
|
|
echo "[entrypoint] /work at $(git -C /work rev-parse --short HEAD) on $(git -C /work rev-parse --abbrev-ref HEAD)"
|
|
else
|
|
echo "[entrypoint] clone FAILED -- the agent has no repository" >&2
|
|
fi
|
|
rm -rf "$_tmp"
|
|
else
|
|
echo "[entrypoint] /work at $(git -C /work rev-parse --short HEAD) on $(git -C /work rev-parse --abbrev-ref HEAD)"
|
|
fi
|
|
|
|
# The shared exchange, for transient files that must not enter git history.
|
|
mkdir -p /exchange/files 2>/dev/null || true
|
|
|
|
# ── Claude Code ──────────────────────────────────────────────────────────────
|
|
# Without this the loop prompt is handed to `exec` as a command, and the whole
|
|
# markdown file is tried as a filename: exit 126, "File name too long".
|
|
if [ "${SYLPH_AUTONOMOUS:-0}" = "1" ]; then
|
|
# Drop the image's default CMD first, or `claude` is handed the literal string
|
|
# "bash" as its prompt and answers a question nobody asked.
|
|
if [ "$#" -eq 1 ] && [ "$1" = "bash" ]; then
|
|
set --
|
|
fi
|
|
# ── Resume across a restart ────────────────────────────────────────────────
|
|
#
|
|
# The container restarts automatically now, and a restart that opens a BLANK
|
|
# session throws away everything the agent knew.
|
|
#
|
|
# 🔴 RESUME BY SESSION ID, NOT BY `--continue`. Measured on the decoder
|
|
# 2026-09-01: `--continue` resolves through `~/.claude.json`'s per-project
|
|
# `history` / `lastSessionId`, and those are written at a GRACEFUL SHUTDOWN --
|
|
# mid-session the live file has both as `None`. A container that is OOM-killed
|
|
# or `docker rm -f`ed never writes them, which is exactly the case this exists
|
|
# for, so `--continue` answered "No conversation found to continue" with the
|
|
# transcripts sitting in the volume beside it.
|
|
#
|
|
# The TRANSCRIPTS are durable and named by session id. Claude Code has not
|
|
# started yet here, so the newest is the previous run's.
|
|
#
|
|
# The /loop prompt is still passed so the loop is RE-ARMED rather than merely
|
|
# restored -- a resumed conversation with no wake-up scheduled answers once
|
|
# and stops, which looks like resuming and is not.
|
|
SYLPH_STAMP="$HOME/.claude/.sylph-last-start"
|
|
SYLPH_RESUME=0
|
|
SYLPH_SESSION=""
|
|
SYLPH_PROJ="$HOME/.claude/projects/$(printf '%s' "$PWD" | sed 's#/#-#g')"
|
|
if [ -d "$SYLPH_PROJ" ]; then
|
|
_newest=$(ls -1t "$SYLPH_PROJ"/*.jsonl 2>/dev/null | head -1)
|
|
if [ -n "$_newest" ]; then
|
|
SYLPH_SESSION=$(basename "$_newest" .jsonl)
|
|
SYLPH_RESUME=1
|
|
fi
|
|
fi
|
|
# 🔴 A POISONED TRANSCRIPT MUST NOT CRASH-LOOP. Restarted under two minutes
|
|
# after the last start, we are already looping: continuing back into whatever
|
|
# killed us is the one thing guaranteed not to help.
|
|
if [ "$SYLPH_RESUME" = "1" ] && [ -f "$SYLPH_STAMP" ]; then
|
|
_last=$(cat "$SYLPH_STAMP" 2>/dev/null || echo 0)
|
|
_now=$(date +%s)
|
|
if [ $((_now - _last)) -lt 120 ]; then
|
|
SYLPH_RESUME=0
|
|
echo "[entrypoint] restarted <120s after the last start -- restart loop"
|
|
echo "[entrypoint] suspected; starting FRESH rather than continuing"
|
|
fi
|
|
fi
|
|
mkdir -p "$HOME/.claude" 2>/dev/null || true
|
|
date +%s > "$SYLPH_STAMP" 2>/dev/null || true
|
|
|
|
if [ "$SYLPH_RESUME" = "1" ] && [ "$#" -eq 1 ]; then
|
|
set -- "$1
|
|
|
|
⚠️ YOU WERE RESTARTED, and this session was resumed — your context is intact,
|
|
but the process that was running when it died is gone. Before anything else:
|
|
|
|
1. \`git -C /work status\`. Whatever you had in progress is still in the tree,
|
|
UNCOMMITTED. Commit it and \`push-work\` before starting anything new.
|
|
2. Any build, test, export or Godot run you had going did NOT finish. Do not
|
|
read its absence as a result.
|
|
3. The likeliest cause is an OOM kill — this container is capped at 4 GB.
|
|
\`CARGO_BUILD_JOBS\` is now set for you in the environment; do not raise it,
|
|
and prefer \`build-export\` / \`build-reference-cli\` over a raw
|
|
\`cargo build --release\`, which bypasses the wrapper's job cap. That is what
|
|
killed the decoder's run on 2026-09-01."
|
|
echo "[entrypoint] resuming session ${SYLPH_SESSION%%-*}… with a restart notice"
|
|
fi
|
|
[ "$SYLPH_RESUME" = "1" ] && set -- --resume "$SYLPH_SESSION" "$@"
|
|
|
|
# Remote Control registers the session with the account so the agent can be
|
|
# reached from claude.ai -- the point of a detached run being that nobody is
|
|
# sitting in front of it. The name is passed EXPLICITLY: the flag's value is
|
|
# optional, so a bare --remote-control swallows the /loop prompt after it.
|
|
if [ "${SYLPH_REMOTE:-1}" != "0" ]; then
|
|
set -- --remote-control "${SYLPH_REMOTE_NAME:-sylpheed-port}" "$@"
|
|
echo "[entrypoint] Remote Control as '${SYLPH_REMOTE_NAME:-sylpheed-port}'"
|
|
fi
|
|
# claude-autonomous wraps `claude --dangerously-skip-permissions` in a pty and
|
|
# answers the one-time first-run gates. The Bypass Permissions disclaimer has
|
|
# no config key that skips it, so unattended it hangs forever.
|
|
set -- claude-autonomous "$@"
|
|
echo "[entrypoint] starting Claude Code in $(pwd)"
|
|
fi
|
|
|
|
exec "$@"
|