Files
Sylpheed/tools/ppc-manual/generator/decoder_scraper.py
sim 7f8a81b8f8
All checks were successful
CI / Native — linux (pull_request) Successful in 2h2m37s
CI / WASM — Web (pull_request) Successful in 30m13s
CI / Formatting (pull_request) Successful in 1m23s
docs(ppc-manual): quote Canary and our own decoder, not the retired xenia-rs
The generator had not been able to run correctly since the manual moved into
`tools/ppc-manual/`: it computed the repository root as `HERE.parent.parent`,
which now names `tools/`, so the XML, Canary's emitters and xenia-rs all stopped
resolving — silently, because both scrapers skipped what they could not find.
Every page's references had been pointing at paths that exist nowhere.

What each source contributed, measured on the 350 pages before this change:

  Operation (pseudocode)  251 pages: fixed boilerplate "derives from the xenia-rs
                          interpreter"; 99 carry real hand-written seeds
  C translation           337 pages: the same kind of boilerplate
  xenia-rs snapshot       336 pages: the interpreter arm, pasted in — the only
                          per-instruction semantics on unseeded pages
  links                   xenia-rs opcode/decoder/interpreter + Canary emitter

Now:

  * semantics come from **Xenia Canary**, the reference emulator, read through
    `git show` at a pinned upstream commit (`origin/canary_experimental`,
    f21ebd49e9). Not our checkout: it carries instrumentation and lacked
    upstream's `mcrf` fix, so it would have published probes and a wrong `mcrf`.
    Each page embeds the emitter (`InstrEmit_<mnem>`), and for the 128 pure
    one-line delegations also the helper that holds the semantics.
  * decode references point at `crates/sylpheed-ppc` — the decoder that
    produces `sylpheed.db` — as in-repo relative links.
  * the boilerplate now says what is true, and the C translation guide maps
    Canary's actual HIR calls, checked against `ppc_hir_builder.h` (including
    that `UpdateCR(n, v)` truncates to 32 bits).
  * `rust_scraper.py` -> `decoder_scraper.py` (interpreter half dropped);
    missing sources are now errors, not empty results.

Verified:

  consistency checks        455 XML entries, 350 families, 598 index keys
  hand-written tails        386/386 byte-identical after regeneration
  xenia-rs in generated     0
  pages with a snapshot     349/350 (was 336) — `dcbi` has no Canary emitter at all
  in-repo decoder links     910/910 resolve to a line holding the identifier
  emitter boundaries        brace counter == column-0 `}` rule on 521/521;
                            preprocessor model unit-tested (#if 0/#else/#elif)
  idempotency               re-run: 0 pages updated, 0 working-tree changes

Hand-written notes (outside the generated regions) are not rewritten here:

  * 110 links into `../../xenia-rs/...` were dead; they now point at the file in
    the archived repository (git.mc02.dev/fabi/xenia-rs @ 8401d4d). Line anchors
    were dropped because the notes predate that commit — 0 of 441 old line
    ranges match it — and a precise-looking wrong anchor is worse than none. The
    link text, which carries the author's line numbers, is unchanged.
  * 140 prose claims about xenia-rs's behaviour remain. 23 are verified to hold
    for Canary too (the 32-bit CR0 truncation, OE left unimplemented); the other
    114 need checking one by one, and some invert — e.g. `divdx` notes a correct
    64-bit CR0 update in xenia-rs where Canary's `UpdateCR` truncates. Left for
    a deliberate pass rather than a blind substitution.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 20:34:34 +02:00

105 lines
4.0 KiB
Python

"""
Locates each instruction in Sylpheed's own PowerPC decoder,
`crates/sylpheed-ppc` — the decoder that produces the disassembly in
`sylpheed.db`, and so the one whose mnemonics a reader of that database sees.
Outputs produced for each mnemonic:
- opcode_line: line in `crates/sylpheed-ppc/src/opcode.rs` where the
`PpcOpcode` variant is declared (1-indexed)
- decoder_line: line in `crates/sylpheed-ppc/src/decoder.rs` where the
variant is produced from raw bits
This used to scrape `xenia-rs/crates/xenia-cpu/src/`, including an
interpreter-arm snapshot. `xenia-rs` is retired and archived, and the decoder
half was lifted into `sylpheed-ppc` unchanged in shape — the same
`pub enum PpcOpcode` and `PpcOpcode::<name>` producers — so only the path and
the interpreter half changed. Semantics now come from Canary; see
`cxx_scraper.py`.
🔴 Missing sources are an error, not an empty index. The previous version
returned `[]` for a file it could not find, so once the manual moved and the
path stopped resolving, every page silently lost its references.
"""
from __future__ import annotations
from dataclasses import dataclass
from pathlib import Path
import re
DECODER_CRATE = Path("crates") / "sylpheed-ppc" / "src"
@dataclass
class DecoderRef:
mnem: str
opcode_line: int | None = None
decoder_line: int | None = None
def _ident(mnem: str) -> str:
"""XML mnemonic -> `PpcOpcode` variant: `.` is not legal in a Rust
identifier, so `addic.` is `addicx` (the same rule Canary uses)."""
return mnem.replace(".", "x")
class DecoderScraper:
def __init__(self, repo_root: Path):
self.src = repo_root / DECODER_CRATE
self._opcode_lines = self._read_lines(self.src / "opcode.rs")
self._decoder_lines = self._read_lines(self.src / "decoder.rs")
self._opcode_index = self._index_opcode_enum()
self._decoder_index = self._index_decoder()
if not self._opcode_index or not self._decoder_index:
raise SystemExit(f"no PpcOpcode variants/producers found under {self.src}")
@staticmethod
def _read_lines(path: Path) -> list[str]:
if not path.is_file():
raise SystemExit(f"decoder source missing: {path}")
return path.read_text(encoding="utf-8").splitlines()
def _index_opcode_enum(self) -> dict[str, int]:
"""Map identifier -> 1-indexed line inside `pub enum PpcOpcode { ... }`
(several identifiers may share a line)."""
idx: dict[str, int] = {}
token = re.compile(r"\b([A-Za-z_][A-Za-z0-9_]*)\b")
in_enum = False
for i, line in enumerate(self._opcode_lines, start=1):
if "pub enum PpcOpcode" in line:
in_enum = True
continue
if not in_enum:
continue
if line.startswith("}"):
break
code = line.strip().split("//", 1)[0]
for m in token.finditer(code):
idx.setdefault(m.group(1), i)
return idx
def _index_decoder(self) -> dict[str, int]:
"""Map identifier -> 1-indexed line of its FIRST `PpcOpcode::<name>`
occurrence, i.e. where the decoder produces it."""
idx: dict[str, int] = {}
pat = re.compile(r"PpcOpcode::([A-Za-z_][A-Za-z0-9_]*)")
for i, line in enumerate(self._decoder_lines, start=1):
for m in pat.finditer(line):
idx.setdefault(m.group(1), i)
return idx
def lookup(self, mnem: str) -> DecoderRef:
ident = _ident(mnem)
return DecoderRef(mnem=mnem,
opcode_line=self._opcode_index.get(ident),
decoder_line=self._decoder_index.get(ident))
if __name__ == "__main__":
root = Path(__file__).resolve().parents[3]
s = DecoderScraper(root)
for m in ("addcx", "addic.", "lwz", "bclrx", "mfspr", "stvx", "vaddfp",
"vaddfp128", "faddx", "mcrf"):
r = s.lookup(m)
print(f"{m:12s} opcode@{r.opcode_line} decoder@{r.decoder_line}")