6263b82ea89ccb1e28ca3114862dc7bd8e31eb3e
1090 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6263b82ea8 |
docs: record the third softening, which was authored dirty
The two instances already in "Checks that were kind once" were correct when written and decayed. The third was wrong on its first commit, and it arrived by a different route: the check and the tree's failure to pass it land in the same change, so the softening writes itself. Concretely — the Clippy step had never run (no component in the toolchain), and the tree is not clippy-clean, so fixing the step and turning it red are the same commit. The first draft paired the fix with `continue-on-error: true` and a comment promising removal once the debt was paid: an expiry date nobody set, in the shape #12's closing line had already ruled out for rustfmt. Reverted on reading it. Adds the distinction, a table separating decay from dirty authorship, and an earlier tell than the mechanical test: If you are writing the softening in the same commit as the check, the thing you want is an issue, not a flag. The mechanical test is unchanged and still correct; this only catches the same failure sooner, at the keyboard rather than at review. Refs #12, #13 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01McNbzUeq1KRBWs4G6X2YVj |
||
|
|
a3d99adaa6 |
ci: install the clippy component the Clippy step needs
`dtolnay/rust-toolchain@stable` installs a minimal profile. The `native`
job named no components, so every run that reached the Clippy step died
on
error: 'cargo-clippy' is not installed for the toolchain
'stable-aarch64-unknown-linux-gnu'
before clippy read a line of source. That is not a lint result; the step
had never run. The `fmt` job below always named `components: rustfmt`
correctly — this one never did.
Two lines of behaviour change. The rest is the comment explaining why the
step is left gating on `-D warnings` rather than softened: the workspace
is not clippy-clean (run 203's build alone emits ~13 rustc warnings that
`-D warnings` promotes to errors), and `continue-on-error` cannot tell
"debt not yet paid" from "debt paid". That debt is scoped in #13, the way
the rustfmt debt is in #12.
Run 203 is what made this visible. With the aarch64 fix in
|
||
| c457320210 |
ci: build for the machine that exists, on the runner that exists
This workflow has never once gone green on this instance: 23 runs cancelled, 2 waiting, zero successes. Not a regression -- it has been decorative since it was written, because it describes GitHub's hosted fleet and runs on one self-hosted aarch64 Pi advertising ["ubuntu-latest","ubuntu-24.04", "ubuntu-22.04"]. Two failures, both configuration rather than code: `windows-latest` and `macos-latest` match no runner label, so those jobs sit in WAITING for ever and the RUN never reaches a terminal state. A pull request's checks therefore never resolve either way -- not red, just never finished, which is worse than red because a red check tells you something. Removed: a second architecture here needs a second runner, not a second matrix row. `--target x86_64-unknown-linux-gnu` on an aarch64 host makes every build a cross-compile, and `wayland-sys`'s build script dies on it with "pkg-config has not been configured to support cross-compilation". Dropped; the native job now builds for its host. NOT touched, deliberately: the WASM and Formatting jobs still fail, on real code state rather than on configuration -- `getrandom` needs the `wasm_js` backend for wasm32-unknown-unknown, and `cargo fmt --check` reports a ~13,000 line diff across the tree. Editing those two into passing is precisely the leniency with an expiry date nobody sets that PROTOCOL.md now forbids. They are issues, not workflow lines. (One latent defect noted while reading: `jetli/trunk-action` fetches trunk-x86_64-unknown-linux-gnu onto this aarch64 host. It has never been reached because the WASM check fails first, and it will bite the moment that is fixed.) Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01McNbzUeq1KRBWs4G6X2YVj |
|||
|
|
25bfa1553f |
protocol: findings before citing code, and checks that were kind once
Some checks failed
CI / Native — ubuntu-latest (pull_request) Failing after 10m46s
CI / WASM — Web (pull_request) Failing after 9m2s
CI / Formatting (pull_request) Failing after 46s
CI / Native — macos-latest (pull_request) Has been cancelled
CI / Native — windows-latest (pull_request) Has been cancelled
Two rules that look unrelated and are one failure, plus the change that makes
the second enforceable.
1. A FINDING REACHES `main` BEFORE THE CODE THAT CITES IT. A citation resolving
only on a peer branch is dead the moment it merges. Not hypothetical: 495
decoder and 366 port commits sit off `main`, and `port/scripts/boot.gd`
already cites two docs/re pages present on neither its own branch nor main.
2. A CHECK MAY ONLY SOFTEN AGAINST A CONDITION IT CAN TEST -- the Pi agent's
wording, and better than mine, because it is applicable while writing rather
than a call to be vigilant. The mechanical form:
Can this branch tell the difference between "not yet" and "no longer"?
`gitea-protect --verify` printed ⚪ "not a collaborator (yet)" and continued,
so the only instrument checking Write-not-Admin could not report that gate
being REMOVED. `check-citations` reported peer citations instead of failing
them, because under the old topology that was unfixable from the container.
Both were correct AND kind when written; neither recorded that the kindness
had a scope. Nobody edits these into being wrong -- the world moves and the
allowance stays, which is why they survive review. The smell is leniency with
an expiry date nobody set; the fix is the testable-condition rule.
check-citations gains `--for-merge`, which turns the peer class into a failure.
A flag rather than a new default because BOTH readings are still live: mid-work
on a topic branch the peer class really is unfixable noise. What the old code
could not express is where the code is GOING, and that is a condition the caller
can state. Measured on this tree: 19 citations resolve only on a peer branch --
which is the size of the #7-depends-on-#8 edge, not the 2 I had counted in
boot.gd.
The selftest gains that third class, because a flag whose classification is
unexercised is the shape this rule exists to catch. Controlled: emptying
PEER_REFS makes the peer case collapse into "nowhere" and the selftest reports
🔴 BROKEN, rc=2.
⚠️ Pre-existing and NOT from this change: the default run already exits 1 on 4
citations of `export/...` paths. Those are the generated tree, gitignored by
design, and main's copy of the tool fails identically. The CITE regex treats
`export/` as a repo prefix. Reported, not fixed -- it is the port's file and its
call whether the regex or the citations are wrong.
|
||
| 7fdcd69434 |
tools: a missing collaborator is a failure, not a blank
--verify's collaborator loop printed ⚪ and continued on 404 without touching `ok`, so the one instrument that checks Phase 1.2 could not report Phase 1.2 being undone. An agent removed from the repository read as "nothing to say" rather than as a gate that is no longer there. It has never fired: Gitea answers that endpoint with permission "read" for a non-collaborator rather than 404, so the case was caught by the role test two lines down. Correct outcome, wrong reason -- the same shape as the check that passed on an instance with no rule at all, and not worth keeping because the luck has held so far. Found by the port agent reading the file rather than running it, which is the only way this one was ever going to surface. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01McNbzUeq1KRBWs4G6X2YVj |
|||
|
|
87932e4017 |
docs: the status block said nothing existed while nine issues were live
Phases 1-4 and 6 are done on the instance. This file still opened with "Nothing
exists on the instance: no agent users, no API tokens, no labels, no milestones,
no branch protection" -- every clause of which was false by the time the merge
that carried it landed.
Replaced with a table of measured state, and each row says what was MEASURED
rather than what was run:
* protection is verified behaviourally -- a real push to main refused with
`pre-receive hook declined`, as the repository owner -- not read off a
settings page. That distinction is the whole subject of this file.
* the tokens are probed: right identity, 403 on branch_protections for both
agents, so the Write-not-Admin carve-out is demonstrated and not asserted.
* the labels are 11 because the instance holds 11.
And a standing note that this block is the part most likely to be wrong, with
what to believe instead: `gitea-protect --verify` and the issue list MEASURE,
this block REMEMBERS. A remembered status is a cache with no invalidation, which
is the same failure as a 1,227-line BLOCKED.md and as the two documents this
runbook was split across an hour ago.
|
||
|
|
bfc6ec4cac | Merge branch 'pi/gate-limit' into agents/gitea-mcp | ||
|
|
a008836ea0 |
docs: the tool creates 11 labels, not 12 -- I counted its own definition
Caught by the Pi agent against the live instance after Phase 4 ran. The tool
creates 5 state/*, 2 agent/*, 4 kind/* = 11.
Where the 12 came from is worth a line, because it is a shape that recurs:
$ grep -c '^mklabel' tools/gitea-setup
12
$ grep -n '^mklabel' tools/gitea-setup | grep -v ':mklabel "'
74:mklabel() { # name colour description
I counted the function DEFINITION as a call. A measurement taken one token away
from the thing being measured -- the same shape as reading protection off a
settings page and reachability off a DNS record, which is now three today. The
version that cannot make this mistake is counting what the instance holds, and
that is what found it.
|
||
| 057bbd438c |
agents: name what branch protection does not gate, and stop the tool contradicting it
Two things that read as protection while being none. Phase 2's rule binds everyone who reaches Gitea through the API or the web, and does not bind anyone with `gitea admin` in the container -- which includes the supervising agent that created the agent accounts and minted their tokens. From that shell the rule is editable and an admin token is one command away. That is the boundary of what the phase buys, not a hole to plug there, and the document read as though the gate were universal. Phases 1 and 2 gate the two CONTAINERISED agents, whose design assumption is that policy lives where they cannot reach it; a supervisor with a host shell is not in that set. And `gitea-setup` finished by telling the reader to go and build a Gitea project board by hand, four sections after the doc explains that a board is a second copy of the state to hand-sync and is precisely the failure that produced a 1,227-line BLOCKED.md. A tool instructing you to do the thing its own documentation argues against is the drift this whole surface exists to end. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01McNbzUeq1KRBWs4G6X2YVj |
|||
|
|
17c0e3f2ab |
docs: fold the page's revisions into the file, and separate wrong from unchecked
The runbook existed as two documents -- a published page and this file -- with
no mechanism keeping them equal, only an intention to remember. Two versions was
the predicted outcome of that, not an accident on top of it. This is the fold,
and the rule that follows it: THIS FILE IS THE SOURCE, the page is derived from
it. When something is urgent enough to push to the page first, it lands here in
the same turn, not "shortly after".
Four things the file did not carry:
* YOUR OWN PUSHES TO main STOP. `enable_push: false` compiles to CanUserPush,
which returns false with no bypass for admins or the owner -- quoted from
the source. Three commits went in by direct push the day this was written,
so the first notice would have been mid-task. Now a check step.
* the token files' MACHINES, which the table had lost.
* do NOT add `write:repository` to the `fabi` token. That scope IS a push
credential. Written down because that advice was given, in chat, by the
author of this file.
* Gitea 1.25.5 confirmed from the desktop too, not just the Pi.
And one thing deliberately NOT folded in: the page said the desktop's outbound
HTTP was blocked, and that is false. `python3 -c 'urllib...'` returns
200 {"version":"1.25.5"} from this box. What is refused here is `curl`, by a
local permission prompt -- which I read as a network constraint and then
published as one. The Phase 3 locations stand; the reason given for them did not.
The "not verified" section now separates WRONG from UNCHECKED. Four entries are
wrong -- requiring an approval does not close the gate, the check could not have
caught that, the token scope, the reachability -- and the pattern in all four is
identical: a property inferred from something ADJACENT to it (protection from a
settings page, reachability from a DNS record) instead of tested directly. That
is the frozen-splash failure, committed in the document about avoiding it. The
first two were caught by the other agent, which is the argument for the review
gate this file exists to build.
|
||
| 7e5719d457 |
tools: apply and re-check the branch protection rule, rather than clicking it
Phase 2 as a file. Six settings where two are load-bearing and both were missing from the first draft is the shape of thing that gets mis-clicked at 1am, so it goes through the API: what was applied is readable in a diff, and `--verify` can re-check it later instead of it being checked once. --verify states its expectations INDEPENDENTLY of what the apply path sends. A check derived from "whatever we posted" cannot fail -- it re-derives the expectation from the thing under test, which is the same instrument-shaped failure as a check that passes on an instance with no rule at all. It also asserts both agents are still Write and not Admin, because an agent promoted to Admin can edit the rule and then merge, so a green rule proves nothing on its own. That is the `gitea-verify` card from "Still to build"; what is left of it is only putting it on a timer. `block_admin_merge_override` stays false on purpose, and the reasoning is in the file: approvals are whitelisted to `fabi`, and Gitea will not let `fabi` approve a `fabi` PR -- so with the override blocked, a human-authored PR could never reach one approval and could never merge at all. The override is not a hole in the agent gate because the agents are Write, not Admin. Phase 1.2 pays for that; this is where it is spent. Reads the repository-scoped credential that already exists on the agent box (~/.sylph-git-credentials) rather than the issue-only ~/.sylph-gitea-api-token, which every branch-protection endpoint refuses. That keeps the setup needing no new credential, and keeps push rights on one machine. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01McNbzUeq1KRBWs4G6X2YVj |
|||
| 034e98eeb0 |
docker: give each agent its own Gitea hands, and close the cross-approval hole
Phase 5 of docs/agents/GITEA-SETUP.md, plus a correction to Phase 2 that the runbook could not have known it needed. gitea-mcp v1.7.0 goes into both images, pinned by the sha256 the release publishes and smoke-tested with `--version` at build time, so a bad pin fails the build instead of the agent. Each entrypoint registers it at user scope for that container's own identity, remove-then-add so a restart is idempotent. The token is passed BY PATH. `-e GITEA_ACCESS_TOKEN=$(cat …)` would write it in cleartext into ~/.claude.json, which every session in the container reads; GITEA_ACCESS_TOKEN_FILE is new in the pinned version and leaves the secret in its read-only mount. Verified against the binary's own --help, not assumed. The tool filter stops being an experiment. The names are in the release README: each agent gets issues, notifications, labels, milestones and pull requests, and NOT `pull_request_review_write`. That one matters because separate identities open a hole the runbook did not name: Gitea refuses to let an author approve their own pull request, and does nothing about sylph-decoder approving sylph-port's. Two agents could satisfy `required_approvals = 1` between themselves and then merge, since branch protection blocks pushes to main and never blocked merges. Withholding the tool is defence in depth; the controls are in branch protection, and both docs now say so: approvals whitelisted to the human so an agent's approval does not count, merges whitelisted to the human so an approved PR is still merged by a person. Phase 2's check gains the step that actually tests it -- approve the throwaway PR yourself, then confirm the agent STILL has no merge button. Without that step, the check passes on an instance where the agents can merge each other's work. Also settles two entries on the runbook's own "not verified" list: the tool filter names, and the Gitea version (1.25.5, whose API schema carries enable_merge_whitelist and enable_approvals_whitelist under those names). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01McNbzUeq1KRBWs4G6X2YVj |
|||
|
|
59649824f6 |
agents: the ordered runbook for standing the Gitea surface up
Some checks failed
WORKFLOW-gitea.md said what the working surface is and why. It did not say how, in what order, or how to know a step worked -- so it was a destination with no route. This is the route. Seven phases, each with a check, each marked 👤 human or 🤖 me: 1 identities two agent users, Write NOT Admin 2 protection main behind a PR + 1 approval -- BEFORE tokens exist 3 tokens three principals, three tokens, three files 4 structure labels and bundles, and deliberately NO Kanban board 5 MCP gitea-mcp v1.7.0, per-agent identity, user scope not .mcp.json 6 items migrate the live findings only -- not 1,227 historical lines 7 restart and verify the three things that must be true Phase 1 leads because it is not hygiene: Gitea does not let a PR's author approve it, so while an agent IS `fabi` either the human cannot approve its work or it can approve its own. The review gate does not exist until the agents are distinct people. (It also fixes 495 commits of agent work attributed to the human's email.) Phase 2's check is a real push and a real PR, not a reading of the settings page. The reason protection lives in the server rather than in a brief is that it should not depend on good behaviour -- so verifying it should not either. Phase 5's install facts are checked, not remembered: gitea-mcp v1.7.0, `gitea-mcp_Linux_x86_64.tar.gz`, `-t stdio -H <host>`, `GITEA_ACCESS_TOKEN`. The `--tools` filter is flagged as an EXPERIMENT that might exclude the merge tool as defence in depth -- explicitly not a substitute for phase 2. Ends with what is still to build (propose-work, an attachment uploader, gitea-verify, the wiki landing page) and a "what I have not verified" section: the approve-your-own-PR behaviour, the --tools names, the Projects API, and the Gitea version -- the API was unreachable from my sandbox three times running. |
||
|
|
cd3a81af31 |
agents: rewrite the briefs for the Gitea workflow
Some checks failed
The two loop files ARE the prompts -- `sylph-port` and `sylph-decoder` read them
off the host at launch -- so the workflow change had to land here or it would
not reach the agents at all.
PROTOCOL.md gains four sections:
* Work items -- issues, milestones as bundles, the state labels, and that
`state/blocked` uses DEPENDENCY EDGES, never prose. A prose blocker is what
let a 1,227-line BLOCKED.md go stale.
* Messages -- an ask is a `kind/ask` issue, not a SendMessage. With the part
that matters: 🔴 NOTHING PUSHES. Notifications are polled, at the top of
every iteration, and therefore an agent must NEVER wait on an ask -- set the
edge, take the next item. The channel this replaces dropped 21 consecutive
messages to a stale session id and reported success each time.
* Pull requests -- one item per branch, `Closes #N`, and you may not merge
your own. Branch protection enforces it; the rule is written down so the
agent knows it, not so it depends on the agent.
* Each iteration, in order -- notifications, sync, one unit, hand over, stop.
Also: evidence a human must look at now attaches to its issue, and a blunt
"never commit game content, under any directory name" with the 545 MB that
prompted it.
The two briefs shrink 697 -> 298 lines. They had accreted five dated focus
blocks between them -- sole-focus orders, F1-F6 queues, one-off "merge this
branch on your first iteration" instructions -- which is a queue, and a queue
belongs in the tracker. What is KEPT is what outlives its bug:
* ask of any check, what would this still report if the feature were absent?
Three instruments passed a splash that never animated.
* the instrument must sit at or above the thing that can break -- the
InputEventAction / input-map miss.
* R1, and grep REFUTED.md before proposing.
* the .pe is primary and the database is somebody's analysis of it.
* the oracle is the real game in Canary, not any renderer of ours.
⚠️ NOT YET TRUE when this lands: the agents have no Gitea users, no API tokens
and no MCP server, so the issue tooling these briefs assume does not exist yet.
The agents are stopped. Setting that up is the prerequisite for restarting them.
|
||
|
|
c3758e3850 |
port: land the play-tested work, and only that
Takes the port branch up to |
||
|
|
1f138b3db4 |
agents: move the working surface to Gitea -- issues, PRs, and where files live
Some checks failed
The human wants to direct this project from a web UI rather than chat or Remote
Control, so Gitea becomes the working surface. No new store: adding a second
copy of the truth is this project's defining failure mode, and Gitea already
holds the code. Its first-party MCP server (gitea/gitea-mcp v1.7.0, checked) has
issues, labels, milestones, PRs, attachments and notifications.
ISSUES replace BLOCKED.md. Milestones are bundles the human defines; issues are
items agents propose and the human approves. The state labels end in
`needs-human`, which is the state the whole model turns on and the one no
off-the-shelf tool models -- the market has converged on removing the human.
`blocked` uses Gitea's DEPENDENCY EDGES rather than prose, so "the Port is
blocked on the Decoder answering X" becomes queryable and closes itself.
PULL REQUESTS, the human's proposal, adopted -- and a bigger improvement than it
looks. Today's long-lived auto/* branches have drifted 280 and 373 commits apart,
which is unreviewable by construction. One PR per item makes the human gate
NATIVE rather than a label convention, binds the change to its item, and enforces
the sizing rule: an item too big to review in one sitting was too big to be an
item.
🔴 Agents must not merge their own PRs, and pull_request_write includes merge --
so this goes in BRANCH PROTECTION on main, not in a document asking them not to.
Same principle that fixed the build-jobs cap: policy where the agent cannot reach
it.
WIKI -- the human suggested it for RE findings, and that half is declined with
reasons. A finding's value is that it sits beside its evidence, versioned with
the code that consumes it; the wiki is a separate git repo, so a decode
correction and the exporter change depending on it could never be one reviewable
PR. And wiki edits bypass review: the REFUTED.md reclassification changed the
file both agents read to decide what not to try, and as a wiki edit it would have
been an unreviewed mutation of shared ground truth. The wiki takes human-facing
orientation instead -- runbook, navigation, container notes, and a landing page,
which closes the real gap that there is no view of what is happening except
container logs.
FILES: three needs, three homes. Agent-to-agent transient stays in /exchange.
Evidence a HUMAN must look at attaches to the issue it belongs to -- it travels
with the item and cannot be orphaned from the claim. Evidence a finding cites
stays in git. Note the MCP exposes attachment_read only; upload needs a direct
REST call.
tools/gitea-setup creates the labels and bundles, idempotently, with --dry-run.
Blocked on a token with write:issue -- the push credential is write:repository
and every issue endpoint refuses it, checked rather than assumed.
|
||
|
|
3a1721abe7 |
docker: stop the wrapper typing into live sessions, and support per-agent logins
Both agents stopped, and the decoder diagnosed it itself:
"I received '2' and '1' but I don't have a pending question those would
answer -- I was in the middle of setting up the /loop cron job."
claude-autonomous matched the BARE SUBSTRINGS 'Choose', 'trust' and 'accept' to
answer Claude Code's one-time first-run gates. The /loop prompt is echoed into
the terminal, and that day's briefs contain 'accepted as-is' and 'least
trustworthy' -- so expect matched the agent's OWN INSTRUCTIONS and typed 2\r and
1\r into a running session, which then sat waiting for a human to explain them.
The old comment argued a multi-word pattern 'never matches' because the gate
text wraps. True of a literal string, false of a whitespace-tolerant regex, which
is what these now are: \s+ spans the wrap, and the terminal is 200 columns wide.
Measured, old against new, against the real brief text and a real gate:
{accept} brief 0 gate 1 (case-sensitive; briefs say 'accepted')
{Yes,\s*I\s+accept} brief 0 gate 1
{trust} brief 1 <- the trigger
{Do\s+you\s+trust\s+the\s+files} brief 0
Two defences, because one is not enough for something that can type: patterns
prose cannot match, and gates skipped ENTIRELY on resume (SYLPH_SKIP_GATES) --
a resumed session cannot show a first-run gate, so there is nothing to answer
and everything to lose. Timeout cut 90s -> 25s for the same reason.
Also: SYLPH_OWN_LOGIN. Remote Control stopped registering under the long-lived
token, and the likely reason is scope -- `claude auth login` requests
user:sessions:claude_code and the token's auth status reports no email, org or
subscription. A per-agent `claude auth login` restores Remote Control AND avoids
the rotation collision, because each agent holds its own grant rather than a copy
of one. The flag stops the entrypoint seeding the host's credentials over it.
|
||
|
|
e53d687755 |
docker: support a long-lived Claude token, and stop the seeding fighting it
Some checks failed
The rotating OAuth credential file is why the agents kept parking, and a long-lived token removes the failure by construction instead of recovering from it after the fact. MEASURED 2026-09-04. ~/.claude/.credentials.json holds a refresh token that ROTATES ON USE. Seeding both containers from the host left three clients holding one token; the first to refresh invalidated the other two, and on the failed refresh Claude Code CLEARS the stored tokens -- writes empty strings, keeps the metadata, and parks at "Login expired". decoder credentials emptied 13:04:28 decoder last transcript 13:04:29 <- one second later The emptying and the park are the same event, which is why it never self-heals: not a stale token a retry could fix, but no token at all, with no browser in the container to complete /login. A hollow file passes every "does it exist" check -- 508 B healthy against 280 B emptied -- which is how three separate diagnoses missed it. And recovery re-armed the bug: after re-seeding, host and decoder held the IDENTICAL refresh token hash. `claude setup-token` issues a long-lived token against the same Claude subscription. Checked, not assumed: `claude auth login` defaults to --claudeai and it is `--console` that means Console/API billing, so this is not the separate API bill. `CLAUDE_CODE_OAUTH_TOKEN` is recognised by the installed binary. Passed as an ENVIRONMENT VARIABLE, both halves of the failure are gone: nothing rotates, so peers cannot invalidate each other, and there is no file for Claude Code to empty on a failure. Both launchers read $HOME/.sylph-claude-token if present -- same pattern as SYLPH_GIT_CREDENTIALS -- and both entrypoints skip OAuth seeding entirely when the variable is set, because copying the rotating file in would re-create the exact collision the token exists to remove. Inert until the file exists. Without it, nothing changes. Also worth recording for the preflight work: `claude auth status` prints JSON with loggedIn/authMethod/subscriptionType. That is a far better SessionStart assertion than checking a file exists, and it would have caught this on the first iteration rather than the third incident. |
||
|
|
f173382fd3 |
docker: the expect wrapper swallowed both the signal and the exit status
Some checks failed
A tooling review predicted a PID-1 signal problem from two symptoms we could not
explain: `OOMKilled: true` with **ExitCode 0**, and `--continue` failing to find
a conversation that plainly existed. Traced it, and the prediction was right --
though the culprit is not PID 1, it is one level below.
The path is tini (PID 1) -> entrypoint.sh (exec'd) -> expect -> spawn -> claude
`spawn` CANNOT be an exec: expect has to stay alive to drive the pty. So expect
is the process Docker signals, and everything depends on it passing things on.
It did neither, in two lines:
1. NO SIGNAL FORWARDING, no trap of any kind. `docker stop` sent SIGTERM to
expect, which died and took the pty with it. Claude Code never got a SIGTERM,
so it never ran SessionEnd hooks and never wrote lastSessionId/history --
which are written ONLY at a graceful shutdown. That is the entire reason
`claude --continue` answered "No conversation found to continue" with 33 MB of
transcripts in the volume beside it, and why we resume by scraping a session
id off a transcript filename.
2. `eof { exit }` RETURNED 0 FOR EVERY DEATH. A bare `exit` in expect is exit
ZERO. When the OOM-killer took the child, expect saw EOF and reported a clean
exit. `OOMKilled: true` with `ExitCode 0` was never Docker being odd -- it was
this line. It also meant `--restart on-failure` would read a memory kill as
success, which is why the policy had to be `unless-stopped`.
Fixed and MEASURED, old against new, in a container:
child exits 7 old -> 0 (the bug) new -> 7
SIGTERM to wrapper old -> 143, child's trap NEVER RAN
new -> 42, child trapped and cleaned up
Same file in both images; they were byte-identical, so the port copy takes the
same change.
Consequences worth stating: a kill now reports 137 rather than 0, so exit codes
mean what they say; `docker stop` gives Claude Code a real SIGTERM, so it runs
SessionEnd and writes the session index -- which may make the transcript-filename
resume unnecessary. That is not assumed here: the resume path stays as it is
until it is verified redundant.
|
||
|
|
fee2e4278a |
agents: one item only -- the title's animation timing -- and split work into human-checkable units
Some checks failed
Two new findings from the human, both about WHEN a title animation starts, and both handed over rather than guessed: F5 Does (A) SNAP the title to finished, or ACCELERATE it? The human says they cannot tell and is right that they cannot -- a three-frame acceleration and a one-frame cut look identical to an eye. Two routes that should agree: a per-frame capture (acceleration shows intermediate alphas, a cut shows none) and the code (assigning a target time and raising a rate multiplier are different instructions). Their "looks more like a snap on multiple attempts" is recorded as a PRIOR, not a result. F6 The title's sweeping white glow -- ptloop01/ptloop02, the blue PCB-like lines -- starts only when the plate appears in the real game, and starts earlier in the port. A lead from the exported declaration, mine and unverified: those elements are keyed at t = 0, 70, 100, 238, 250 while the plate reaches full alpha at 236, with pteff02 keyed at exactly 236 and ptlogo_back2eff and ptcopyright at 238. 236-238 is a synchronisation point in the declared data and a human just reported a behaviour change there. Flagged AGAINST itself too: 238...250 looks equally like an exit ramp -- ptcopyright uses that shape and starts nothing -- and the sweep lives in a nested .rat leaf with its own timeline. F6 bears on clock: "shared" and on F4: if a title element does not move until the plate arrives, either the declared data says so and our keyframe reading is wrong, or something at the plate's arrival STARTS it, which is a mechanism nobody has proposed. And the process change, which is the human's and outlives this item: "attacking the 'whole' mission was too big for them to handle. Split the given missions and tasks into even smaller tasks which they can tackle and give to a human for feedback." PROTOCOL.md gains "Work in units a human can check in a minute". A milestone is not a unit of work, it is a bag of them. A unit is right-sized when it ends in something a person can judge in under a minute WITHOUT READING ANYTHING, and each one states its question, what the human looks at, and what it does NOT cover. Do one, hand it over, stop -- an unverified fix under a second change makes a regression two-variable. The evidence for the rule is this week: the splash sat through a whole milestone and took one day once scoped to "does it animate?". The bar is a HUMAN check, not a green tool -- three instruments passed a frozen screen. |
||
|
|
6438316f24 |
agents: correct "both clocks" -- there is ONE, and F4 tests whether it is right
Some checks failed
I wrote "whether the game snaps both clocks forward" into yesterday's F4 and the human asked which clocks. There are none: authored/flow.json sets `clock: "shared"`, so the title's two composited builds -- build 4 the artwork (finishes t~=118) and build 2/3 the plate (full alpha t=236) -- run on ONE clock started together. Left standing, that phrasing sends an agent hunting for a second clock this corpus says does not exist. Corrected in both briefs and in the playtest page, marked as a correction rather than silently edited. And the question is better than I first framed it. `clock: "shared"` is AUTHORED, and the port's own plate-arrival-halves.md calls it "not falsified... not confirmed to better than ~20 % either", with an unresolved anchor disagreement inside one binary: the reconciliation picked t=118 while settle_time() returns 160 and the boot prints "settles at t=160". So F4 is a TEST OF THAT PREMISE, and the discriminator is observable -- press (A) early, while the wordmark is still building in, and watch the ARTWORK rather than the plate: advances the shared clock -> the artwork SNAPS to finished only forces the plate -> the artwork KEEPS ANIMATING its build-in Both briefs now say to answer F4 before building on `shared`, and tell the port not to choose what "jump" means. |
||
|
|
18620e99aa |
agents: P5's gate is MET, and four findings from the same walk
Some checks failed
"Menu walk and navigation is fine. Video skips too. Extras open. New Game shows new game intro video." -- 2026-09-02 P5 is done. Its gate was "a human clicks through it", the retro said it had been waiting on that and not on code for the whole milestone, and it has happened. PORT-MISSION.md updated. The NEW GAME gap is accepted as-is. Four findings, three of them the Decoder's: F1 THE MENU REPEATS ON A HELD DIRECTION AND OURS DOES NOT. One step per deflection was authored as the conservative choice because nobody knew; a human has now watched the real game and it repeats, "at a medium pace... slow enough to see which item is selected". That settles the existence half of H1 against us. The RATE is still unmeasured and must not be guessed -- the description bounds it and supplies no number. Decoder measures initial delay and repeat interval as frame counts; the port implements the mechanism and waits for the numbers. F2 THE SFX ARE TOO LOUD BECAUSE THERE IS NO MIX AT ALL. Measured: confirm -17.7 dB mean / -0.0 dB peak, 3 dB hotter in mean than the music and 6.4 dB above move. No volume or gain value exists anywhere in export/ or authored/, so every clip plays at unity on one bus. Decoder: is per-cue or per-bus gain on the disc -- the cue table is the obvious place and cue 1103 is already decoded. Port: gains at PLAYBACK as data, and explicitly NOT normalisation in the exporter, which destroys the relationship between clips and cannot be undone by a modder. F3 SOMETHING IS MISSING ON THE TITLE SCREEN. The export carries one music file and the port plays nothing on the title. Which cue does the title play, and is there a sting on the plate or on accept? A negative needs a positive control: find the menu's cue by the same method first. F4 (A) SKIPS FORWARD THROUGH THE BOOT AND WE IMPLEMENT TWO OF THREE PRESSES. In the game: skip video, reveal plate immediately, accept plate. The middle one is missing here. Whether the game snaps both clocks forward or only reveals the plate is a question, not a detail -- and it is a cheap second route to the plate-arrival question, since a press that skips to the plate says where the game thinks the plate belongs. H3, the plate delay, is ACCEPTED -- "feels the same... sufficient". Left unattributed rather than closed green. |
||
|
|
0ba7542547 |
agents: the logo splashes are DONE -- the human cannot tell them from the game
Some checks failed
"Looks good! Cannot notice any obvious difference from the actual game.
Mark logos as done." -- 2026-09-02
Not "the check passes": a person compared the port against the real game and
could not tell them apart. That is the oracle, and it is the strongest result
this port has produced. The sole-focus order is lifted; both agents return to
their milestones.
The fix was one word -- pose_at ASSIGNED the settle instant instead of clamping
to it, so every query returned the settled pose whatever the clock said. The
same line manufactured the false green: the capture harness shoots after two
frames, so it was photographing t~=2 units, which looked settled only because
everything looked settled. The 0.01 % agreement that closed H2 was measured
through the accident. One bug produced the defect AND the evidence of its
absence.
Verified here before it went to the human, by film rather than by claim:
motion 16.4 % -> 27.7 %, distinct luma states 26 -> 43, the publisher ramp 6
steps -> 13 in one continuous run, and the developer splash's interrupting
0.50 s freeze gone. The publisher trajectory rises to a peak and settles back --
the crossfade signature.
The port then closed a gap motion-census names in its own header ("a wrong ramp
that moves every frame passes here") with a shape check pre-registered from the
disc, measured off a film, on a non-overlapped strip, in ratios so the texture
divides out: rise:last declared 1.20, measured 1.20 exact.
Kept as the standing lesson, because it is the fourth instance: an instrument
that sits below the thing under test cannot see it fail. Ask of any new check
what it would still report if the feature were entirely absent.
Explicitly NOT claimed: P5's gate is "a human clicks through it" and nobody has
said the milestone is met. The briefs say so, and say not to record it on the
human's behalf.
The decoder's end-to-end pipeline work returns to normal priority rather than
being dropped -- it is what decides whether the port's 60 units/s matches the
game. The ramp is now right in SHAPE and unverified in DURATION.
|
||
|
|
3cc3400a96 |
agents: the splash does not animate, and three instruments could not see it
Some checks failed
A human on a GPU at ~140 fps: "the logos just switch, there is no animation at
all." Measured from a real boot with --film at 0.05 s, then per-frame change:
splash moves 1.30 s of 7.95 s = 16.4 %
publisher splash 0.30 s of motion, then 3.20 s FROZEN
developer splash 0.35 s + 0.25 s, then 2.40 s FROZEN
distinct luma states in 7.95 s 26
A 45-unit build-in cannot be drawn in 26 states, and a fade does not hold one
picture for 3.20 s. The frame counter says 24.8 fps achieved; both are true --
the port is DRAWING 25 times a second and CHANGING almost never.
🔴 Why every check passed, which matters more than the bug:
frozen sweep drives the clock BY HAND -- proves the renderer can draw
pose N, never that the poses are drawn in sequence
settled compare 0.01 % against the capture -- a screen frozen 84 % of the
time matches a settled reference PERFECTLY, that is what
frozen means
achieved fps counts frames DRAWN -- the same pixels 25x/s scores
identically to animating
Every one measured throughput or a pose. None measured CHANGE. Same shape as
InputEventAction bypassing the input map: the instrument sat below the thing
that was broken, so the break could not appear in it.
tools/motion-census closes the class. It measures change and nothing else, and
its --selftest asserts it separates a fade (97.4 % moving) from a switch (2.6 %)
from a frozen film (0.0 %) -- a detector that cannot tell those apart would
report the same green line on all three.
Both briefs: this is the SOLE focus. The port reproduces before changing
anything and gates every fix on a film rather than a still. The decoder maps the
whole pipeline end to end -- disc bytes, the game's per-frame update (does it
interpolate between keyframes or hold?), what is submitted per frame, and what
Canary does to it before a capture records it -- delivered as a SERIES, not a
settled value.
The port should also record the refutation against itself: H2 reads ANSWERED on
the strength of the frozen sweep. The mechanism half stands, the blur is a baked
companion texture. The behaviour half does not.
|
||
|
|
47207a1474 |
port: give the port container a GPU path -- it never had one
Reported as "the port has low FPS". Godot 4 renders through Vulkan and this launcher passed nothing through, so it fell back to lavapipe: software Vulkan, correct and slow. The decoder's launcher has had this block for a long time; the container that actually runs a renderer was the one without it. Same three cases as the decoder, including the part worth repeating: passing /dev/dri alone does NOT work for NVIDIA -- Mesa cannot drive the card and the proprietary userspace lives outside the image. It needs the container toolkit. The NOTE now prints the full repo-add sequence, because the package is not in Ubuntu's default repos and `apt install nvidia-container-toolkit` on its own fails with 'no installation candidate' -- which reads like the package is wrong rather than the source being missing. |
||
|
|
e1749c83e1 |
docker: auto-restart, and resume the session the agent was actually in
Some checks failed
The decoder died mid-task and it took four separate findings to explain, each of which read as something else: 1. OOM-KILLED, REPORTED AS A CLEAN EXIT. `OOMKilled: true` with **ExitCode 0**. So `--restart on-failure` would treat a memory kill as a successful finish and leave the agent down -- the policy has to be `unless-stopped`. 2. THE JOB CAP WAS SET AND THEN REMOVED THREE LINES LATER. build-reborn has always exported CARGO_BUILD_JOBS, but a raw `cargo test --release -p sylpheed-formats` never reaches the wrapper. Adding `-e CARGO_BUILD_JOBS` to the launcher did not help either: the entrypoint recomputes and exports over it unconditionally. An explicit value now wins, and says so in the log. 3. THE MEMORY CONSTANT WAS WRONG. `mem_gib * 2 / 3` assumes ~1.5 GB per job; release rustc on this workspace needs ~2 GB, and 4 jobs in 6 GB is what died. Divisor is now 2. 4. `--continue` CANNOT RESUME AN ABRUPT DEATH, which is the only kind we get. It resolves through ~/.claude.json's per-project `history`/`lastSessionId`, and MEASURED mid-session both are None -- they are written at a graceful shutdown. A killed container never writes them, so `--continue` answered "No conversation found to continue" with 33 MB of transcripts in the volume beside it. Persisting .claude.json did not help, because the fields were never populated in the first place; that attempt is removed rather than left in looking useful. The TRANSCRIPTS are durable and named by session id, so the entrypoint reads the id off the newest one for its cwd and passes `--resume <id>`. Verified on both agents: each reattached to its exact prior session and appended to the same file rather than opening a new one. The /loop prompt is still passed alongside `--resume`, so the loop is RE-ARMED rather than merely restored -- a resumed conversation with no wake-up scheduled answers once and stops, which looks like resuming and is not. Restarting into the same death is guarded at the other end: a start less than 120 s after the previous one begins FRESH instead of continuing back into whatever killed it. That fired correctly during this work. On resume the agent is told it was restarted, that its in-progress work is uncommitted in the tree, that any build or capture it had running did not finish and its absence is not a result, and which wrapper to prefer over a raw release build. |
||
|
|
1af103d9b9 |
agents: point each brief at its human branch, to merge on the first iteration
Some checks failed
Both are pushed. The decoder's carries the R1 register reclassification and tools/stale-instrument; the port's carries the two input fixes, verify-input and BLOCKED H1-H3. Each branches from that agent's own tip, so it is a fast-forward on the line they are already on -- and the port must merge before touching input or it will re-derive a fix that is already asserted. |
||
|
|
aad3fb382e |
agents: the splashes exactly, and stop photographing a moving thing
Some checks failed
A human played the port on real hardware for the first time (2026-09-01) and
found four things. Two were port defects, fixed. Two are open and are now both
agents' focus: the PRESS (A) plate arrives late, and the splash fade/blur is
weaker than the game's.
Their verdict on method is the reason this is a brief change and not a ticket:
"the agents were essentially guessing and trying to copy what one would see,
but while they did get close it still is not quite right"
Close-but-not-right is the signature of reproducing APPEARANCE instead of
deriving MECHANISM. So the Decoder's focus block asks, in order: is there a
post-process pass at all, what is it, where do its parameters come from -- and
only then what curve. Both routes, dynamic (GPU state, shader constants, render
targets; add logging to Canary, it is theirs read-write) and static (.pe, the
DB, the paks), with each fact labelled by which produced it.
TEMPORAL-VERIFICATION.md is the other half, and it generalises past the
splashes. We have been photographing the game at time t, and t is never the
same twice: emulator speed varies with host load, Canary presents at ~28.1 fps,
the capture path costs a variable 0.1-10.8 s, and a long-lived x11grab stream
degrades and then freezes. The register already carries FOUR refutations of
exactly this shape. The replacement rule: record a film, not a photograph;
align by CONTENT, not by clock, and report the lag as a measurement rather than
minimising it away; prefer ordering, counts, durations and shape over any value
at a wall-clock instant; anchor on an event; report achieved fps against
requested fps.
Also into both briefs: the input set. The port had no joypad binding for (A) or
(B) and nobody noticed for a whole milestone, because --script sends
InputEventAction, which BYPASSES the input map -- so every check asserted the
code below the map and nothing about the map. The Decoder is asked to DECODE
the full set the game reads rather than discover it by pressing buttons; the
Port is told input is verified at the device level or not at all.
And both briefs now point at the R1 register reclassification, because two of
the ten re-opened entries land on this focus: "the declared keyframe timeline
reproduces the captured splash" is 🟡 our-reader, and the rest() pair is open
in BOTH directions -- while the two splashes are the only screens that reach
that fallback.
|
||
|
|
1b1a4dfcd3 |
containers: an expired token could never be replaced
Some checks failed
Credentials were seeded only when the container's copy was MISSING. So when a session expired, the file still existed, the copy was skipped, and restarting changed nothing -- the one recovery path a human has, re-logging in on the host, could not reach the containers at all. Now re-seeds whenever the host's copy is newer. Newer-wins rather than always-copy, because a container refreshes its own token mid-run and that copy may legitimately be the fresher of the two. Found when both sessions expired: host credentials at 16:30, containers holding 14:20 and 14:24. |
||
|
|
d1685d67c9 |
viewer: show where a cutscene's voice actually is, and let you hear it
Some checks failed
The Cutscenes window printed the voice token as text and offered no way to play it, which left the most confusing thing on the disc invisible. The movie voices are one continuous XMA stream chunked into VOICE_*.slb entries whose boundaries do NOT match the cutscene cues, so the bank named after a movie need not hold that movie's audio. Measured, on the retail disc: ADV region 433930240..437044592 inside VOICE_ADV.slb name honest S00A region 452798464..455499120 inside VOICE_S00A.slb name honest RT01A region 437044592..437345648 inside VOICE_ADV.slb NAME LIES RT01A's voice sits in bytes belonging to the entry named after the intro movie. A viewer that played the name-matched bank would be confidently wrong for exactly the cutscenes where it matters, and would look right on the two that are easiest to check. So the window now shows BOTH locations -- the named bank with its byte range, and the resolved region -- and states plainly whether the name is honest, highlighting it when it is not. Play routes through the movie form of RequestAudio, which resolves the region rather than reading the bank. Static data only: sound.pak and tables.pak, both on the disc. |
||
|
|
f1b87e47b6 |
decoder: tell it about the reference assets, and that the DB can be wrong
Some checks failed
The mounts landed but the agent could not learn of them: I documented them in CONTAINER-NOTES.md, which the decoder's prompt does not list, and then restarted the container -- so a fresh session with no memory of the exchange had a 586 MB database and a decompressed image sitting unmentioned in its filesystem. Now in the PROMPT itself, not only in a document, because the prompt is the one thing a new session is guaranteed to read. CONTAINER-NOTES.md is also added to its reading list. And the caveat that matters more than the asset. The .pe is PRIMARY -- the bytes the console executed. The database is somebody's ANALYSIS of them, produced by a disassembler that had to guess, and it is wrong in the ways disassemblers are wrong: misdecoded mnemonics where data was read as code, function boundaries short or long or merged or split, coverage missing entirely for code reached only by indirect dispatch, and names that are derived rather than symbols. So a finding resting on a database row is not established until the bytes agree: read the same address out of the .pe and check. Where they disagree the image wins, and the disagreement is itself worth recording, because it tells the next reader which parts of the database to distrust. A fast index into 9.2 MB of machine code, not a source of truth. |
||
|
|
72b10e7d03 |
decoder: mount the disassembly DB and the flat VA image
Some checks failed
The decoder had neither, and reported the gap precisely: four scripts in this repo READ /work/xenia-rs/sylpheed.db and nothing produces it, so the whole static PPC route was consumers with the producer missing. Both exist on the host and are now mounted read-only: the 586 MB database (25 481 functions, 851 classes with RTTI, EH tables, imports, 1.8M indirect-dispatch candidates) and the decompressed image. The image is the more useful of the two. It is a FLAT VA DUMP -- file offset = VA - 0x82000000 -- so reading a known address needs no XEX decrypt, no LZX, and no booted emulator. The decoder had independently recovered the same bytes by dumping /dev/shm/xenia_memory_* and validating against the GamePart table, which is good work and a sound method, but it noted itself that needing a running emulator is a bad dependency for something the entire static corpus rests on. It does not need one. Also recorded that an earlier claim the .pe was STALE was tested and refuted, so nobody re-litigates it, and that instructions.raw is an INT rather than hex. Written down as reference material, explicitly NOT a deliverable: they are read-only, they come from outside the repository, and a fresh checkout elsewhere has neither. Reimplementing the producer belongs in sylpheed-formats, and until it exists every static finding rests on an artefact this project cannot rebuild. |
||
|
|
2021eee47d |
agents: merge main at the start of every iteration
Some checks failed
Both agents read the protocol, their mission and the shared tooling from their OWN checkout, and both work on topic branches -- so without an explicit sync they follow whichever version of the rules existed when the branch started. Found concretely: tools/audio-capture and two protocol revisions were on main while the decoder worked for hours from a branch that had neither. The port had merged on its own initiative and did have them, which is exactly the kind of divergence nobody notices until the two disagree about what the rules say. |
||
|
|
a8d2491366 |
audio: actually install the capture path I kept deferring
Some checks failed
The audio work was three parts and I shipped two. The transcode-fidelity method and the pinned 5.1 downmix landed; the null sink -- the only one that answers "what does the GAME play" -- I deferred to "the next natural rebuild window" and then rebuilt both images four times without doing it. pulseaudio-utils is now in both, with tools/audio-capture wrapping it: a null sink is a real device as far as an application is concerned, so Canary and Godot open it normally and parec records what they emit. This unblocks the decoder's Q8. The cue-to-event bindings are currently a name match against the authors' own identifiers -- a plausible guess, not a measurement -- and capturing what the game plays on a menu move converts them. `audio-capture run` reports the peak level and warns when the capture is silent, because silence is the failure that looks like success: a WAV of exactly the right duration, full of zeroes, because the application opened a different sink. A duration check alone passes it, which is how a confident wrong number gets made. |
||
|
|
20b3c74b2c |
agents: they never spoke, the decoder lost the disc, and both shared one state dir
Some checks failed
Three defects, all mine, found by checking instead of assuming. **They never exchanged a word.** SendMessage=0, ListAgents=0 across both new sessions. PROTOCOL.md specified in detail what a message MAY and MAY NOT do and never said how to send one or that the other agent was addressable -- they knew that last time only because the human told them directly, and rebuilding with fresh volumes wiped it. Policy without mechanism is prose. Now documented with the two addresses, a worked example, and an instruction to introduce themselves on the first iteration rather than waiting to have a question. **The decoder lost the disc and the ISO.** They used to arrive inside the project mount and silently stopped when /work became a clone. Silently is the word: the disc-gated tests SELF-SKIP without SYLPHEED_DISC and report green, so a whole test suite would have passed while measuring nothing. Both are now mounted explicitly, the ISO at a stable path so run-canary does not depend on host directory names. **Both agents shared one Claude state directory.** They share the host's ~/.claude, and once both working directories became /work they resolved to the same projects/-work/ -- two supposedly independent agents writing to one place, which undoes the point of separate checkouts. Each now has its own volume, seeded once from the host with credentials only, so a token refresh writes locally and neither can corrupt the host's auth. Also widened the pacing rule. It banned ScheduleWakeup by name; the decoder then scheduled itself an hourly cron job -- not harmful, but the same instinct that ended a run yesterday, through a door I had left open. Now: no self-scheduling by any route. Mount audit after the changes: shared and intentional are the exchange volume and the read-only credential seed. Everything else -- repo, Claude state, cargo, target, canary, disc, ISO -- is per agent or one-sided. |
||
|
|
01a3505b1e |
containers: fix volume ownership and make the clone guard survive interruption
Some checks failed
Two bugs, both mine, both found by starting the thing.
**Volume mount points must exist AND be owned by the agent before USER agent.**
Docker seeds a named volume from whatever the image has at that path, ownership
included, and creates a ROOT-OWNED directory when the path is absent. Either way
the agent cannot write, and the failure surfaced far from its cause: "clone
FAILED", with no permission error anywhere in sight. The port's own Dockerfile
already carried a comment explaining this trap, which I then walked into for
/work and /exchange.
**The clone guard checked for a .git directory, not a usable HEAD.** A clone
interrupted partway -- the container was removed while one ran -- leaves a .git
with no commits, and a presence check then skips the retry forever and hands the
agent an empty repository that looks like a checkout. It now verifies HEAD, and
clones via a temp directory so a partial result never lands in /work at all.
Also: the port launcher's path defaults still assumed the old repo root, so it
mounted no disc; and the stale /reborn notice is gone now that there is one
repository.
Verified running: both agents cloned
|
||
|
|
06676d3dc0 |
containers: each agent clones the monorepo into its own volume
Some checks failed
The last structural fix for the collision class that has bitten three times. Both containers now clone the repository into their OWN named volume instead of bind-mounting a human's working tree, so an agent's local git config cannot capture a human's commits, a credential helper cannot leak a container-only path onto the host, and a `git add -A` cannot sweep another party's in-flight files. Cloned once at startup and never auto-pulled: pulling under a running agent moves files out from under whatever it is mid-edit, which is the same bug again. Accepted knowingly: Claude Code keys per-project memory off the working directory, so moving off the host path starts that memory empty. The corpus in docs/ is the memory that matters and it travels with the clone. Other changes: * docker/agent -> docker/decoder; the launcher is sylph-decoder. Roles, not "the agent", now that there is more than one. * /reborn is gone -- one repository now, so the port reads HANDOFF from its own checkout rather than through a live read-only mount of someone else's tree. * Canary mounts separately at /canary; it stays a fork tracking upstream. * A shared `sylpheed-exchange` volume at /exchange, with tools/ on PATH so `share` is available in both. * The decoder's credential file gets the .host-copy treatment the port already had -- `credential.helper=store` rewrites by rename-over-target, which is EBUSY on a bind mount and reports a fatal that is not one. * Budget split deliberately: decoder 5 cpu / 6 GB, port 3 / 4, leaving room for the planned Referee. "Half the host" was right when there was one agent. Prompts move to docs/agents/ and are rewritten around the protocol: the oracle is the running game, dynamic RE stays with the decoder, each iteration must attempt to refute one claim of the other, and neither may verify its way out of its own role. |
||
|
|
a8815f2826 |
agents: the team protocol, the share tool, and a player's-eye navigation doc
Some checks failed
**navigation.md rewritten from the player's chair.** It was written from the inside out -- GamePart ids, pak names, sprite names -- which is how WE find things, not what the game shows anyone. Now it describes what is on screen, what you press and what happens, with internals as footnotes. Most rows are open on purpose: it exists to be filled in by playing, and the in-game tutorials are the resource for the flight half. **tools/share** gives transient files provenance without giving them history. Three kinds of thing were travelling down one channel with opposite needs: code and decoded knowledge want permanence, cited evidence wants permanence, and "look at this PNG" wants no history at all. The third kind bloats a repository forever; passing it by message is worse, because the receiver gets bytes with no idea which build produced them. `share put` records who, when, what, the sender's commit, and whether their tree was dirty -- because a capture taken from a modified tree cannot be reproduced from the sha, and the receiver deserves to know that before building an argument on it. **docs/agents/PROTOCOL.md** is the contract. The parts that matter: Dynamic RE stays with the Decoder -- most of what is open is behavioural and cannot be answered from the file. What the planned Referee adds is different: bias enters at what you CHOOSE to capture, so a corpus captured to a fixed protocol by someone with no hypothesis is worth more than one captured to settle an argument. A message may point, ask, prioritise and challenge. It may not change scope, redefine ground truth, or carry a finding instead of writing it down -- including a message claiming to relay the human, because a relayed instruction has no evidence attached and this project has watched a wrong belief travel further and faster than its correction. Adversarial duty is explicit: every iteration, try to refute one claim of another agent and record the attempt either way. Run your own instrument through a control first. Disagreements go to the human with both positions, not to whoever is more certain. And no agent may verify its way out of its own role: the Port has no oracle, the Decoder builds nothing, the Referee interprets nothing. |
||
|
|
65cefa74c3 |
monorepo: one repository for the decoders, the port and the corpus
Some checks failed
Merges the Godot port into the reverse-engineering repository, preserving both
histories -- 1019 commits of corpus plus the port's 31, brought in by subtree
merge and then moved into place so git can follow each file across the rename.
The reason is not tidiness. The two-repo split forced the exporter to depend on
the decoders by pinned revision, and that created a whole class of failure that
now disappears: a sha reachable only from a topic branch, orphaned by a
squash-merge, breaking a fresh checkout silently at build time. It also forced a
live read-only mount of one agent's working tree into another's container, which
is why a contract file could move mid-iteration. With a path dependency, a
decoder change and the exporter change it requires land in the same commit or
not at all.
Canary stays separate: it is a fork tracking upstream.
New structure for the long term:
docs/game/ how the game is NAVIGATED -- menus, modals, prompts, alerts,
and in-game flight. Written so nobody rediscovers it. Mostly
open questions on purpose; the in-game tutorials are the
resource for the flight half.
docs/port/MODDING.md
modding as a constraint on the exporter TODAY, not a later
feature: one logical asset in one file (the disc splits nearly
everything, and resolving that is the exporter's job), names a
person recognises, PNG/OGG/OGV/JSON only, base-and-overrides so
re-exporting is always safe, provenance in every file.
data/base + data/mods
generated tree and drop-in overrides, both gitignored
exchange/ transient inter-agent files, deliberately outside history
docs/agents/ the team protocol
Both the README and the navigation doc lead with the correction that cost the
most: the oracle is the real game under Xenia Canary. Reborn's renderer is a
hypothesis under test, it has been wrong, and treating it as ground truth
propagated into three documents and both agents before a human caught it.
Scripted modding stays possible without being built: no screen name is hardcoded
in GDScript and there is no native code in port/, which is what Godot Mod Loader
needs to be able to substitute behaviour later.
|
||
|
|
8c33f86a20 |
merge the Godot port's history into the monorepo
Brought in with a subtree merge rather than a copy, so the port's 31 commits survive as history rather than arriving as one anonymous import. Landed under godot-import/ and moved into the final layout in the next commit, which keeps git's rename detection able to follow each file across the move. |
||
|
|
ff0937fa8a |
port: pin the 5.1 downmix matrix explicitly -- human decision
The RE agent escalated this rather than picking, correctly: folding centre into L/R changes how dialogue sits against the music, which is an aesthetic judgement about the game and not a container detail. Decided: explicit ITU fold, centre at -3 dB, recorded in the manifest with the rest of the transcode command. Pinned rather than defaulted because a default is a decision nobody made -- invisible in the output and free to change between ffmpeg versions. |
||
|
|
8d7373b16c |
video: state the 5.1 downmix instead of inheriting it, and write atomically
A real defect in the P4 output, found by the human on ADV.wmv and widened by the RE agent to the whole disc: the disc ships 28 movies in 5.1 WMA Pro (every cutscene, INCLUDING both movies this port needs) and 69 already in stereo. A bare `-ac 2` therefore does two different things and records neither -- stereo passes through, and 5.1 is folded by ffmpeg's DEFAULT matrix. How loudly centre-channel dialogue sits against the music is a content decision, and it was being made by accident and could move under an ffmpeg upgrade. Now stated: ITU-R BS.775, LFE dropped, normalised by 1/(1+2*sqrt(1/2)) = 0.4142. It appears in the recorded command, so the manifest determines the output. MEASURED rather than chosen by taste, and the measurement is the interesting part: the explicit matrix and ffmpeg's inherited default differ by a residual of -91 dB -- about one LSB at 16-bit -- with peak and mean agreeing to 0.1 dB. So ffmpeg's default IS this matrix, and the audio does not change; what changes is that the manifest now says which matrix. The UNnormalised textbook form was measured too and clips at 0.0 dBFS, which is why the scaling is there. The filter is applied only to 6-channel sources, probed per file with ffprobe, so a stereo source is never run through a matrix referencing channels it lacks. Also: encode to a temp name and rename on success. ffprobe read a mid-write .ogv as 33 s against a 137 s source -- no error, no warning, the exact shape of catastrophic truncation. The filesystem is shared with the RE agent, so that is a race, not an edge case, and a half-written file must never be visible under its final name. tools/verify-video-audio answers the second of the three questions docs/AUDIO-VERIFICATION.md separates: does GODOT route the audio. An AudioEffectRecord on the Master bus writes Godot's own mixed output to a WAV from a headless run, so "no sound card" was never the obstacle I claimed. It deliberately checks non-silence and level only -- a difference-signal RMS against the source is inconclusive without cross-correlation alignment and an agreed downmix, and would produce a confident wrong number. |
||
|
|
700094bef1 |
FORMAT v3: rotation, and the focus record -- the ring the port could not reach
Pin bumped to the TAG formats-pin-2026-08-29 (
|
||
|
|
c0a5725934 |
port: pin by tag, and how to verify audio with no sound card
**The pin.** Pin a TAG, never a bare sha. A sha reachable only from an auto/* branch is orphaned when that branch is deleted or -- worse -- squash-merged, because squash creates new commits: main looks like it contains the work while the pin becomes unreachable and this project stops building for a fresh checkout, silently, at their build. formats-pin-2026-08-29 exists for the current state. Also says plainly why NOT to float to a branch, which was the tempting fix: Cargo resolves a git dependency once and writes the sha into Cargo.lock, so floating gives staleness you cannot see instead of staleness you can read. push-work now pushes --follow-tags so annotated tags travel with the branch. **Audio.** docs/AUDIO-VERIFICATION.md separates three questions that were being asked as one: is the transcode faithful (no engine, no device -- a file-vs-file difference measurement), does Godot route it (AudioEffectRecord on the Master bus writes a WAV from a headless run), and what does the GAME play (a PulseAudio null sink, which needs a rebuild). It leads with the three ways the fidelity measurement lies, because all three were hit on the first attempt and each produces a confident wrong number rather than an error: unaligned subtraction, mismatched channel layouts, and probing a file another process is still writing. The 5.1 disc fact deliberately is NOT copied here -- it lives in the RE corpus at docs/re/structures/movie-audio-channels.md and is linked, so there is one copy to keep true rather than two that drift. Same reason HANDOFF is a summary with links. The downmix itself stays flagged as an unmade decision, not quietly resolved. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
26b5afeaa3 |
agent: tag the decoder states the port pins, and stop leaking credential config
Two fixes to push-work and the policy that goes with them.
**Tags.** The port's exporter depends on sylpheed-formats BY REVISION, so a
commit of ours is part of its build -- and the commit it pinned lived on one
auto/* branch and nowhere else. Deleting that branch orphans it; squash-merging
it is worse, because squash creates NEW commits, so main appears to contain the
work while the pinned sha becomes unreachable and the port stops building for a
fresh checkout. Silently, at their build, long after the breakage.
push-work now pushes with --follow-tags, which publishes annotated tags
reachable from the pushed commits, and MISSION.md says to tag whatever the port
needs. formats-pin-2026-08-29 at
|
||
|
|
3491d30066 |
re(media): the disc ships movies in TWO audio profiles, and 28 of them are 5.1
Probed all 97 movies. 28 are wmapro 48 kHz 6-channel 5.1 -- ADV.wmv and every S*.wmv story cutscene; the other 69 are wmav2 48 kHz stereo, every RT*.wmv and hokyu_*.wmv. The split is cinematics vs in-mission radio chatter. Both movies the menu milestone needs, ADV.wmv (the boot/attract intro) and S00A.wmv (the new-game intro), are in the SURROUND group. Why it matters: one ffmpeg command over dat/movie/ produces two different kinds of result and records neither. The 69 stereo files pass through unchanged; the 28 surround files get downmixed 5.1 -> stereo by ffmpeg's DEFAULT matrix, folding centre-channel dialogue into L/R at a weighting nobody chose and which is not stable across ffmpeg versions. That is a content decision inherited by accident, so it should be stated explicitly and recorded beside the command. Credit where due: found by the human while checking a transcode, verified independently here and widened from one file to the whole disc. Two METHOD entries from the same episode, both about measurement rather than format: don't probe a file another process is still writing (a half-written transcode reported 33 s against a 137 s source, no error, nearly a filed bug), and a difference-signal RMS is meaningless before cross-correlation alignment (-34.2 dB against a -25.3 dB source looks like failure and is inconclusive). |
||
|
|
7eeae3006a |
re(ui): the focus ring SPINS, the game draws it, and the leaf owns the f record
Three things, all from parsing ptbtn0Nf.rat as a build. 1. THE RING SPINS. Its two keyframes differ in exactly one field: rotation_deg ramps 0 -> 360 with position, scale, alpha and tint all constant. A spin in place, the same shape as the GP_BUNK example already recorded. 2. THE ORACLE CONFIRMS THE GAME RENDERS IT. In the OPTIONS-focused capture the ring's bright head sits in a completely different angular position from the sprite's own -- caught mid-spin. This is a SECOND independent confirmation that rotation_deg is drawn, now on a different screen and a different element from the ptloop sweeps, and it raises rotation's priority: it is not a title-only concern that sits off-screen at rest, it is the main menu's focus marker. NO ANGLE IS QUOTED. A brightest-region centroid says ~250 deg, but the control refuses that precision -- rotating the sprite by a known 30/90/180/270 and re-measuring gives errors up to 19.8 deg. What survives the error bar is that a <=20 deg error cannot manufacture a ~250 deg displacement. 3. WHICH PLACEMENT WINS -- correcting this page's own earlier caveat, which said to use the leaf only for elements the parent does not declare. Right for a BASE record, wrong for an f record: the parent declares NO element for ptbtn0Nf.rat at all (zero of build 5's 16), so the f record's placement comes from its leaf for BOTH elements, label included. The label's (-7,-7) is load-bearing -- the f sprite is 13px larger per axis and -7 keeps them concentric (535+96/2 = 583 vs 542+83/2 = 583.5). Corroborated against the oracle: the focused-minus-unfocused region is x 505..703, and the leaf predicts a right edge near 707 where the parent reading predicts 714. Also exposes UiBuild::records (name -> (offset, size) of a nested .rat leaf). Nested records were parsed into a PRIVATE map, so a consumer holding a UiBuild could not locate a leaf's bytes at all -- which is exactly what blocked the port from reaching the ring. |
||
|
|
2fbcd59920 |
docs: retract "reference renderer" -- sylpheed-cli is not the oracle
A framing correction from the human, and it runs through everything I have written, so it is a retraction rather than a silent edit. Reborn "was/is just a GUI explorer and extraction CLI for verifying the decoding of the various files. It may very well be wrong." The oracle is the Xenia Canary capture and the game. So verify-screen is a CONSISTENCY check between two decoders that share their assumptions, plus a regression detector -- not a correctness check, and agreement in it is not evidence of correctness. Its header now says so, it calls the CLI the COMPARISON renderer, and DIFFERS means "we moved apart, find out which of us moved". The uncomfortable part, recorded because it is the actual failure mode: this file already contained the sentence "two renderers reading one field through one decoder agreeing is not evidence that the field is right", written after the ptframe1 case -- and I then went on quoting 3/255 against sylpheed-cli as though it meant the port was right. Having the principle written down did not stop me leaning on the agreement. Three times both renderers agreed and both were wrong, each caught only by a capture: pteff05 (menu screens had no background), scale-0 (drawn full size instead of collapsed), rest() (the menu bracket missing). Correctness moves to the captures -- nine of them, indexed at docs/re/captures/ORACLE-CAPTURES.md, covering all five screens in scope. Three cautions travel with them: not gamma-neutral (there is a floor, don't chase it), geometry IS sound (a positional disagreement is real), and each is one moment of a still-animating screen. verify-screen keeps running over all 16 screens every iteration. It is still worth having -- total, cheap, and it catches a divergence introduced on the RE side. It is just not a grade. |
||
|
|
579c8096c1 |
port: P4 -- the intro video plays inside the boot sequence
The exporter transcodes ADV.wmv and S00A.wmv to Ogg Theora and records the exact ffmpeg command in the manifest, per MISSION §6, so a modder who dislikes the quality re-runs one line rather than reverse-engineering what was done. Quality was MEASURED, not judged: SSIM against the decoded source over a 10 s sample is 0.9863 / 0.9896 / 0.9924 at -q:v 6 / 8 / 10, and at 200 % zoom on the reel's hardest case -- fine serif text and soft gradients over near-black, where Theora breaks first -- q8 is indistinguishable. So MISSION §6's permitted FFmpeg-GDExtension fallback is NOT needed and is NOT being proposed. No new runtime dependency. -ac 2 because the source is 6-channel WMA Pro; that downmix is a decision, so it lives in the recorded command rather than in prose. Encoding is cached on a .cmd sidecar holding the command and the source size -- any change to either re-encodes. export/ is still regenerated wholesale; this is derived state validating derived state, not a hand-edit, and without it every re-export pays ~4 minutes to produce a byte-identical file. The player renders INTO the design SubViewport. Parenting it to the Boot node played the movie to the window instead, and every captured frame came out black -- which is worth more than a capture-bug note: a movie outside the 1280x720 design space is outside the coordinate system every screen is expressed in. (A) skips a movie, because Q9 measured that (title at 57 s vs a 193 s baseline). NOT VERIFIED, and stated as such: audible playback. This container has no audio device and Godot falls back to the dummy driver. The Vorbis stream exists, is 2-channel and decodes; whether Godot emits it is unconfirmed. |
||
|
|
d110cf38c7 |
media: expose se_wave_riff -- the menu's SE cues, assembled where the format lives
The port is forbidden from reimplementing media assembly and Static.slb is exactly that case: no RIFF, no seek chunk, no XACT container, just a packed run of whole 2048-byte XMA1 packets, so a wave is defined only by (offset, packet count) and the header has to be synthesized. That step now happens once, in the crate that owns the format, instead of in each consumer. `slb::xma1_wave_riff` wraps raw packets; `media::se_wave_riff` looks the bank up and reads just the packets asked for. Both reuse the existing synth_xma1_fmt / build_riff, which are already byte-identical to what tools/re-capture/ slb_extract_wave.py writes -- so this is exposure, not a second implementation. It reads a TARGETED range rather than the whole bank, and that is load-bearing: Static.slb is the ONE entry of sound.pak's 9 519 whose declared extent runs past the end of the extracted segments -- by exactly 616 768 B -- so reading it whole fails outright on this extraction. Every cue we need is in the first few hundred KB. Recorded rather than worked around silently. Verified as an artifact, not a compile: all three cues decode through ffmpeg to mono 48 kHz PCM at 0.533 / 0.344 / 1.016 s, non-silent (rms 2085 / 2985 / 4327, peaks 29813 / 16973 / 32767). The refusal path is exercised in the same run -- an impossible packet count is rejected rather than returning a short stream, because a truncated XMA decodes to plausible-sounding garbage. Also adds docs/re/captures/ORACLE-CAPTURES.md: an index of the nine canary framebuffer captures already in this repo, and a plain statement that THEY are the reference and `screen render` is not. |