3 Commits

Author SHA1 Message Date
f4d59c5783 re: Stage 02 is lost at ~11 minutes, and the pilot's survival rule is what guarantees it
First session whose deliverable was the mission's ENDING rather than a
measurement (mission_run.sh, 500 s, hull of every entity at 2 Hz). The ACROPOLIS
is untouched to t=170 s then falls at ~53 HP/s with no let-up, reaching zero at
t=640-720 s — so "no mission completed" is not an artifact of the 240 s
time-boxes, and not of the 600 s cap on a blocking tool call. A longer session
would only watch the loss arrive.

Attributing the damage by co-presence, exactly two classes are ever near the
asset: e007 turrets (8483) and e010 bombers (8334). pilot.py treats turrets as
keep-out zones at 2500 units and never as targets — the rule that made it
survive — so roughly half the escort damage comes from the one class it is
designed to avoid. Survival and the objective are in direct conflict and the
pilot resolves it entirely for survival: WARSHIPS 0000, WARPLANES 0009,
REMAINING OB rising 004 -> 008, our hull untouched at 1500/1500 with 120
missiles spent. That is unspent risk budget, not a good run.

Also corrects launch_mission.sh: a harness-tracked BACKGROUND task does not keep
the display alive (lost 11 s in, at the turn boundary) — the
one-blocking-foreground-call rule stands.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 20:57:43 +00:00
530555de9f re: the weapons are the control — sibling-default inheritance is unit-schema-specific, not engine-wide
Four rules from 21 units invites the coincidence objection, so run the identical
sweep against the Weapon/Shell capture, which has COMPLETE coverage (126
records). It finds no sibling rule at all: the one 100%-agreement candidate has a
single distinct value and is really a constant default. Weapon defaults vary per
record exactly as unit defaults do, so "defaults are computed" is general while
"defaults come from a sibling field" is not.

Size_Y <- Size_X survives, and is now checked at the raw-token level rather than
through the sub-record merge: e105, f105 and f101 each declare Size_X/Size_Z/
Size_Radius and no Size_Y, and each reads back its own Size_X at runtime. The
two two-unit hypotheses are demoted to coincidence-not-excluded.

Also records a negative for planning: Stage 01, the only other reachable stage,
adds four uncaptured units that are variants of already-captured ones, so it
would re-measure rather than test the rules.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 20:09:14 +00:00
69b4a2e569 re: a unit field left unset on disc is not a global constant — Size_Y inherits Size_X
Attacks the 21/110 unit-coverage limit from the cheap side: if a defaulted field
always took one runtime value, the captured units would pin it for all 110. Only
6 of 24 confirmed defaulted fields behave that way. The other 18 vary per unit,
so the default is computed.

Asking which OTHER field of the same unit holds that value — counting only
non-zero cases, and checking the two fields sit at different offsets so the
layout solver cannot be aliasing them — gives four rules. Size_Y <- Size_X is
solid: seven unrelated ships (e105 600, e106 300, e108 80, e201 300, f101 400,
f105 700, f106 200) each omit it on disc and each shows its own Size_X live,
while the two fields differ freely when both are on disc. Size_Radius fits both
min(X,Z) and the median of the three axes and cannot yet be separated;
e010_ADAN_Attacker_S is what rules out the simpler Size_X rule. FCSRange and
DefencePoint rest on two independent units each and ship as HYPOTHESIS.

Applied across the disc the rules recover 65 (unit, field) values in units that
have never been visited. Falsification test recorded: load any uncaptured stage
and compare one predicted value.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 19:44:24 +00:00
6 changed files with 234 additions and 2 deletions

View File

@@ -22,7 +22,7 @@ Promote to a prose `structures/…md` file when a format needs behavioural notes
| XBG7 mesh | 🟡/❔ | `sylpheed-formats/src/mesh.rs` + `tests/mesh_disc.rs` ([xbg7](structures/xbg7-mesh.md)) | weapons/props: declaration-driven variable stride (36 models), GPU-confirmed. **Stage containers: 5662 sub-models across 22 stages** via content-anchored grouped pools (`stage_models`). Quantized hero bodies (DeltaSaber `f004`) still declined |
| Capital-ship part placement | 🟡 | `sylpheed-formats/src/ship.rs` (static) + [runtime capture](ship-placement-runtime-capture.md) | hull placement static-exact; external parts approximate statically. **Runtime capture** (Canary F10 → VS-constant WorldView) gives ground truth — validated on `e106` destroyer; not yet baked into the viewer |
| Weapon fields defaulted on disc | ✅ | [runtime struct](structures/weapon-struct-runtime.md) · [DATA SHEET route](weapon-datasheet-runtime.md) | **Solved.** Canary maps guest RAM into `/dev/shm`, so the parsed `Weapon`/`Shell` objects are readable live; their layout is solved against disc ground truth (zero contradictions over 100+ records). All 126 weapons, exact numbers, no story progress needed — [4 393 values](captures/weapon-runtime-fields.csv) the disc does not carry. Supersedes the letter-bucket limit of the DATA SHEET route, which now serves as the independent cross-check |
| Unit (craft/vessel) fields defaulted on disc | ✅/🟡 | [runtime struct](structures/unit-struct-runtime.md) | The parsed `unit\UN_*.tbl` definition object, vtable `0x820af844`, ≥`0x380` bytes, one per unit — **discovered, not assumed** (`unit_discover.py`), and distinguished from the spawned-entity class `0x820af030` by being one-per-ID and byte-constant within a run. Across runs only pointer words move — `--crosscheck` proves **no reported field offset is run-dependent** (two words, `+0x2c8`/`+0x2d0`, are stage-dependent and remain unidentified). 27 fields ✅ (21 units, 7 runs); the `Maneuver` block is **schema declaration order, 4 bytes/field, base `0x9c` with a two-slot gap after `AA_Roll_Min`** (29 anchors, 0 conflicts), which also pins 5 fields *no* disc record ever values. Angles are **radians at runtime, degrees on disc**. Unlike weapons, unit definitions are instantiated **per stage**, so coverage (21/110) grows by visiting missions — [values](captures/unit-runtime-fields.csv) |
| Unit (craft/vessel) fields defaulted on disc | ✅/🟡 | [runtime struct](structures/unit-struct-runtime.md) | The parsed `unit\UN_*.tbl` definition object, vtable `0x820af844`, ≥`0x380` bytes, one per unit — **discovered, not assumed** (`unit_discover.py`), and distinguished from the spawned-entity class `0x820af030` by being one-per-ID and byte-constant within a run. Across runs only pointer words move — `--crosscheck` proves **no reported field offset is run-dependent** (two words, `+0x2c8`/`+0x2d0`, are stage-dependent and remain unidentified). 27 fields ✅ (21 units, 7 runs); the `Maneuver` block is **schema declaration order, 4 bytes/field, base `0x9c` with a two-slot gap after `AA_Roll_Min`** (29 anchors, 0 conflicts), which also pins 5 fields *no* disc record ever values. Angles are **radians at runtime, degrees on disc**. Unlike weapons, unit definitions are instantiated **per stage**, so coverage (21/110) grows by visiting missions — but a defaulted field is **not** a global constant: `Size_Y` provably inherits `Size_X` (7 independent units, 6 distinct values), and three more sibling rules are recorded ❔, recovering 65 values in units never visited — [values](captures/unit-runtime-fields.csv) |
| UI screen layout (`.rat`) | ✅/🟡 | [ui-rat-layout](structures/ui-rat-layout.md) | One pak per UI screen; each RATC = one (context × language) build; every `<name>.t32` sprite has a `<name>.rat` **layout record** (BE u32; 1280×720 design space; scale/tint/X/Y, keyframes for animated elements, `opt ` link to the focused state). **The tutorial PAUSE menu and the title main menu both rebuild pixel-accurately from the disc.** `loop1.rat` (screen-level draw order) not yet decoded |
## Runtime / dynamic-capture technique

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.7 MiB

View File

@@ -0,0 +1,89 @@
# Why Stage 02 is never won — the escort sinks at ~11 minutes (2026-08-10)
**Status: ✅ measured, one 500 s run.** The standing open item since 2026-07-29 was
"no mission completed". This is the first session whose deliverable was the
*ending* rather than a measurement, and it settles why: **the mission is lost
before it can be won, and the pilot's survival policy is what guarantees it.**
Run: `tools/re-capture/mission_run.sh 500 mission01` — boot → Stage 02 in flight →
`pilot.py` (escort-weighted targeting, target commitment, guided missiles) for
500 s, with every entity's hull sampled at 2 Hz and a screenshot every 30 s.
Artifacts at `/sylph-home/re/mission01/` (9 MB `mission.jsonl`, not committed).
## The escort's decay is linear, and it ends the mission
| t (s) | ACROPOLIS hull | % |
|---|---|---|
| 0160 | 25000 | 100 % |
| 180 | 24510 | 98.0 |
| 280 | 20279 | 81.1 |
| 380 | 13799 | 55.2 |
| 480 | 8541 | 34.2 |
| 485 (end) | 8182 | 32.7 |
Untouched until **t ≈ 170 s**, then **≈53 HP/s** with no let-up — so the asset
reaches zero at **t ≈ 640 s**, and the whole-run average rate puts it at 722 s.
Either way the escort is dead at **1012 minutes**, and "the ACROPOLIS is sunk"
is a defeat condition ([mission-escort-state](mission-escort-state.md)).
This also retires a suspicion: the 240 s time-box of earlier runs was *not*
hiding a win, and the ~500 s ceiling of a single blocking tool call is **not**
the binding constraint. A longer session would simply watch the loss arrive.
## What is actually killing it — and the conflict that follows
Attributing damage by co-presence (which hostiles are within 3000 units of the
asset in the sample where its hull drops, damage split evenly among the classes
present — suggestive, not per-shot proof), only **two** classes are ever near it:
| class | samples present | attributed damage |
|---|---|---|
| `UN_e007_ADAN_Turret` | 200 | 8483 |
| `UN_e010_ADAN_Attacker_S` | 194 | 8334 |
Roughly half the damage comes from **turrets** — and `pilot.py` treats turrets as
**keep-out zones at 2500 units, never as targets**. That rule is not arbitrary: a
turret is what shot down every pilot before 2026-07-30, and it is why the craft
now survives. But it means **the policy that keeps the pilot alive also
guarantees the escort dies.** Survival and the objective are in direct conflict,
and the pilot currently resolves it entirely in favour of survival.
The HUD at t≈485 s says the same thing from the game's side:
- `YOU KILLED WARSHIPS` **0000** — not one warship in 500 s, across every run ever;
- `YOU KILLED WARPLANES` **0009** — fighters only;
- `REMAINING OB` **004 → 008** — objectives are being *added* by waves faster than
any are cleared, so the pilot is not touching the objective set at all;
- SHIELD and ARMOR bars full, hull **1500/1500**, 120 missiles spent.
![Stage 02 at t≈485 s](captures/mission01-t485.png)
## The conclusion that matters
The pilot optimises the wrong thing. It maximises survival and fighter kills;
the mission scores **objectives** and **the escort**, and the fighter population
(134 → 92) is close to irrelevant to both. An untouched 1500/1500 hull at the
moment the escort passes 33 % is not a good run — it is **unspent risk budget**.
Concretely, for the next attempt, in priority order:
1. **Turrets near the asset must become targets**, not keep-out zones — accepting
hull damage is the only way to cut ~50 % of the incoming escort damage. The
keep-out rule should be scoped to turrets that are *not* threatening the
asset, rather than applied globally.
2. **Engage warships.** `WARSHIPS 0000` forever means the objective class has
never been attacked; `REMAINING OB` rising is the scoreboard saying so.
3. Re-check whether the escort damage rate actually falls once turrets die —
that is the experiment that tells us whether (1) is sufficient or whether the
bombers need dedicated intercept too.
## Method note, learned the hard way
A harness-tracked **background** task does *not* protect the display: the same
run launched in the background lost Xvfb 11 s in, at the turn boundary
(`skip_intro` exit 3, "DISPLAY LOST"). The comment in `launch_mission.sh` saying
the script may be run as a tracked background task is **wrong**; the
one-blocking-foreground-call rule still stands, which caps a single attempt at
the tool's 600 s timeout. And do not pipe a long run through `tail` — the first
attempt printed nothing because `timeout` killed the pipeline before it flushed;
the on-disk artifacts are what survived.

View File

@@ -272,3 +272,84 @@ take-off, ~12 minutes in mid-combat, and after GAME OVER. 14 objects, the same
is no need to play it, and no need to survive it.
Stages captured so far: `Ttrl` (BASIC CONTROLS), Stage 02.
## A defaulted unit field is not a global constant — some inherit from a sibling
**Confidence: 🟡 for `Size_Y`, ❔ for the rest. Analysis 2026-08-10, offline, from
[`captures/unit-runtime-fields.csv`](../captures/unit-runtime-fields.csv).**
The coverage limit above (21 of 110 units, growing only with story progress) is
worth attacking from the other side first: *if* a field the disc leaves unset
always took the same runtime value, the 21 captured units would pin that default
for all 110 and no further missions would be needed.
**It does not.** Restricting to the 150 values that are both ✅ CONFIRMED and
come from a field the disc leaves defaulted, only 6 of 24 fields have a single
value across every unit that defaults them (`HP`→10, `MassScore`→0,
`MaximumVelocity`→0, `RadarRange`→0, `DestroyMotionTime`→0, `Size_Z`→0.1). The
other 18 take several distinct values — so the default is computed per unit.
Where from? For each defaulted value, ask which *other* field of the same unit
holds exactly that value. Counting only cases where the value is **non-zero**
(otherwise `0 == 0` inflates every pair) and checking that the two fields are at
**different offsets** (so the match is not the layout solver aliasing them):
| defaulted field | takes the value of | support | independent units |
|---|---|---|---|
| `Size_Y` (`0x034`) | `Size_X` (`0x030`) | 9/9 | **7**, 6 distinct values |
| `Size_Radius` (`0x050`) | `min(Size_X, Size_Z)` | 4/4 | 4, 3 distinct values |
| `FCSRange` (`0x2a4`) | `RadarRange` (`0x2a0`) | 4/4 | 2 |
| `DefencePoint` (`0x2bc`) | `AttackVesselPoint` (`0x2b4`) | 6/6 | 2 |
`Size_Y ← Size_X` is the one to trust: seven unrelated ships (`e105` 600,
`e106` 300, `e108` 80, `e201` 300, `f101` 400, `f105` 700, `f106` 200) each omit
`Size_Y` on disc and each shows its own `Size_X` at runtime. When both fields
*are* on disc they differ freely (14 distinct `Size_Y` values against 13 of
`Size_X`), so this is a default rule, not one value stored twice.
`Size_Radius`'s formula is **not yet separable**: `min(Size_X, Size_Z)` and "the
median of the three axes" fit all four units identically. `UN_e010_ADAN_Attacker_S`
is what rules out the simpler `Size_Radius ← Size_X` (X=100, Y=40, Z=50, radius
**50**). The last two rules rest on two independent units each and are ❔ —
recorded so they can be falsified, not relied on.
**Why it matters for the reimplementation:** filling a missing `Size_Y` with `0`
or with a global constant gives the game's largest hulls a wrong lateral extent
(`f105` 700, `e105` 600, `f101` 400 — all defaulted on disc). Applied across the
disc, the rules recover **65 (unit, field) values in units that have never been
visited**: `Size_Y` in 21 of the 21 units that omit it, `Size_Radius` in 22 of 26,
`FCSRange` in 14 of 56, `DefencePoint` in 8 of 60.
### Cross-check against the weapons: this is NOT an engine-wide mechanism
The obvious worry is that four rules from 21 units are coincidence. The
`Weapon`/`Shell` capture is the control: **complete coverage, 126 records**, with
the same "defaulted on disc" classification. Running the identical sweep there
(confirmed rows, non-zero values, offsets required to differ) finds **no sibling
rule at all** — the single 100 %-agreement candidate (`Shell.Length ←
`Shell.Volume`, 5 records) has one distinct value, i.e. it is really the constant
`Length → 10` coinciding with `Volume = 10`. Weapon defaults vary per record just
as unit defaults do (10 of 14 `Weapon` fields, 16 of 17 `Shell` fields), so the
phenomenon is general; the *sibling* explanation is not.
So `Size_Y ← Size_X` is **specific to the unit schema** (plausibly the size block
defaulting its axes), not a property of IDXD default resolution. Two consequences:
the rule cannot be justified by appeal to a general mechanism, and the two
two-unit hypotheses (`FCSRange`, `DefencePoint`) lose the support they would have
borrowed from one — treat them as **coincidence-not-excluded** until a new stage
tests them.
`Size_Y ← Size_X` itself survives this scrutiny, and was re-checked at the raw
token level rather than through the sub-record merge: `UN_e105_ADAN_Cruiser`,
`UN_f105_TCAF_Cruiser` and `UN_f101_TCAF_Acropolis` each declare `Size_X`,
`Size_Z` and `Size_Radius` and **no `Size_Y` at all**, and each reads back its own
`Size_X` (600 / 700 / 400) at runtime.
**How to falsify:** the rules predict a specific number for units in stages not
yet captured. Load any new stage, snapshot, and compare — one disagreement kills
the rule. Note what is *not* a useful test: Stage 01, the only other reachable
stage, adds just four uncaptured units (`e010`/`e106` variants) whose predictions
are the same numbers their already-captured base variants gave, so it would
re-measure rather than test. A real test needs a stage with unfamiliar classes,
i.e. story progress — which is now the *only* thing story progress is needed for
here.

View File

@@ -19,7 +19,9 @@ alive(){ ps -o pid=,stat= -C xenia_canary 2>/dev/null | awk '$2 !~ /^Z/ {print $
# that actually bought was the opposite: a process nothing owns is a process
# nothing keeps alive, and both were being reaped a couple of minutes in — the
# long-standing "Xvfb and the emulator die on their own every few minutes" note.
# Run this whole script as ONE tracked background task and leave Xvfb, openbox
# MEASURED WRONG 2026-08-10: a harness-tracked BACKGROUND task does not protect
# them either — the display was lost 11 s in, at the turn boundary. Run this
# whole script as ONE BLOCKING FOREGROUND call and leave Xvfb, openbox
# and xenia as its children: they then live exactly as long as the session does.
# `nohup` still shields them from a stray HUP; the exit-status wrapper means a
# death is reported with the server's own account instead of being inferred.

60
tools/re-capture/mission_run.sh Executable file
View File

@@ -0,0 +1,60 @@
#!/usr/bin/env bash
# Attempt to COMPLETE a mission and record how it ends.
#
# Every previous flight session was time-boxed to 240 s to measure something
# (escort hull, lethality, ship placement) and none ever reached a mission
# outcome — "no mission completed" has been the standing open item. Unit
# definitions are instantiated per stage, so story progress is the only thing
# that grows unit coverage past 21/110, and that needs a WIN, not a survival.
#
# So this run is deliberately long and its only deliverable is the ENDING:
# screenshots throughout, every entity's hull sampled, and the pilot log kept
# whole (never tail-piped — a frozen log tail is what mission-end looks like
# from outside, and tailing throws away the transition).
#
# Runs as ONE tracked background task with Xvfb/openbox/xenia as plain nohup
# children — see docs/re/session-lifetime notes; do NOT setsid anything.
#
# Usage: mission_run.sh [flight_seconds] [tag]
set -u
export HOME=/sylph-home/re SDL_AUDIODRIVER=dummy DISPLAY=:98
export PYTHONPATH=/sylph-home/.local/lib/python3.12/site-packages
SD="$(cd "$(dirname "$0")" && pwd)"
SECS="${1:-900}"
TAG="${2:-mission}"
SHOTS=/sylph-home/re/shots
OUT="/sylph-home/re/$TAG"
mkdir -p "$SHOTS" "$OUT"
"$SD/launch_mission.sh" fly || { echo "BOOT FAILED"; exit 1; }
python3 "$SD/entities2.py" self 0x130 "$OUT/cfg.json" || { echo "BIND FAILED"; exit 1; }
echo "=== initial entity table ==="
python3 "$SD/mission_state.py" scan "$OUT/cfg.json"
# A screenshot every 30 s for the WHOLE run: the outcome card (MISSION COMPLETE
# / GAME OVER) is on screen only briefly, so sampling must not stop early.
( n=$(( SECS / 30 + 4 ))
for i in $(seq 1 "$n"); do
printf '%s SHOT %03d\n' "$(date +%s)" "$i" >> "$OUT/shots.log"
screenshot "$SHOTS/$TAG-$(printf %03d "$i").png" >/dev/null 2>&1
sleep 30
done ) &
SHOTTER=$!
date +%s > "$OUT/t0"
python3 "$SD/mission_state.py" watch "$OUT/cfg.json" "$SECS" 2 "$OUT/mission.jsonl" \
> "$OUT/mission.log" 2>&1 &
WATCHER=$!
python3 "$SD/pilot.py" "$OUT/cfg.json" "$SECS" > "$OUT/pilot.log" 2>&1
PILOT_RC=$?
wait $WATCHER 2>/dev/null
kill $SHOTTER 2>/dev/null
screenshot "$SHOTS/$TAG-end.png" >/dev/null 2>&1
cp -f "$SHOTS/$TAG-end.png" "$OUT/end.png" 2>/dev/null
echo "PILOT_RC=$PILOT_RC"
echo "--- last 5 pilot lines ---"; tail -5 "$OUT/pilot.log"
echo "--- shots: $(ls "$SHOTS/$TAG-"*.png 2>/dev/null | wc -l) ---"
echo "MISSION RUN DONE ($TAG, ${SECS}s)"