Files
Sylpheed/tools/ppc-manual/memory/lvx.md
sim dedcf37867 docs(ppc-manual): quote Canary and our own decoder, not the retired xenia-rs
The generator had not been able to run correctly since the manual moved into
`tools/ppc-manual/`: it computed the repository root as `HERE.parent.parent`,
which now names `tools/`, so the XML, Canary's emitters and xenia-rs all stopped
resolving — silently, because both scrapers skipped what they could not find.
Every page's references had been pointing at paths that exist nowhere.

What each source contributed, measured on the 350 pages before this change:

  Operation (pseudocode)  251 pages: fixed boilerplate "derives from the xenia-rs
                          interpreter"; 99 carry real hand-written seeds
  C translation           337 pages: the same kind of boilerplate
  xenia-rs snapshot       336 pages: the interpreter arm, pasted in — the only
                          per-instruction semantics on unseeded pages
  links                   xenia-rs opcode/decoder/interpreter + Canary emitter

Now:

  * semantics come from **Xenia Canary**, the reference emulator, read through
    `git show` at a pinned upstream commit (`origin/canary_experimental`,
    f21ebd49e9). Not our checkout: it carries instrumentation and lacked
    upstream's `mcrf` fix, so it would have published probes and a wrong `mcrf`.
    Each page embeds the emitter (`InstrEmit_<mnem>`), and for the 128 pure
    one-line delegations also the helper that holds the semantics.
  * decode references point at `crates/sylpheed-ppc` — the decoder that
    produces `sylpheed.db` — as in-repo relative links.
  * the boilerplate now says what is true, and the C translation guide maps
    Canary's actual HIR calls, checked against `ppc_hir_builder.h` (including
    that `UpdateCR(n, v)` truncates to 32 bits).
  * `rust_scraper.py` -> `decoder_scraper.py` (interpreter half dropped);
    missing sources are now errors, not empty results.

Verified:

  consistency checks        455 XML entries, 350 families, 598 index keys
  hand-written tails        386/386 byte-identical after regeneration
  xenia-rs in generated     0
  pages with a snapshot     349/350 (was 336) — `dcbi` has no Canary emitter at all
  in-repo decoder links     910/910 resolve to a line holding the identifier
  emitter boundaries        brace counter == column-0 `}` rule on 521/521;
                            preprocessor model unit-tested (#if 0/#else/#elif)
  idempotency               re-run: 0 pages updated, 0 working-tree changes

Hand-written notes (outside the generated regions) are not rewritten here:

  * 110 links into `../../xenia-rs/...` were dead; they now point at the file in
    the archived repository (git.mc02.dev/fabi/xenia-rs @ 8401d4d). Line anchors
    were dropped because the notes predate that commit — 0 of 441 old line
    ranges match it — and a precise-looking wrong anchor is worse than none. The
    link text, which carries the author's line numbers, is unchanged.
  * 140 prose claims about xenia-rs's behaviour remain. 23 are verified to hold
    for Canary too (the 32-bit CR0 truncation, OE left unimplemented); the other
    114 need checking one by one, and some invert — e.g. `divdx` notes a correct
    64-bit CR0 update in xenia-rs where Canary's `UpdateCR` truncates. Left for
    a deliberate pass rather than a blind substitution.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 20:34:34 +02:00

7.0 KiB
Raw Permalink Blame History

lvx — Load Vector Indexed

Category: Memory · Form: X · Opcode: 0x7c0000ce

Assembler Mnemonics

Mnemonic XML entry Flags Description
lvx lvx Load Vector Indexed
lvx128 lvx128 Load Vector Indexed 128

Syntax

lvx [VD], [RA0], [RB]
lvx128 [VD], [RA0], [RB]

Encoding

lvx — form X

  • Opcode word: 0x7c0000ce
  • Primary opcode (bits 05): 31
  • Extended opcode: 103
  • Synchronising: no
Bits Field Meaning
05 OPCD primary opcode
610 RT/FRT/VRT destination
1115 RA/FRA/VRA source A
1620 RB/FRB/VRB source B
2130 XO extended opcode (10 bits)
31 Rc record-form flag

lvx128 — form VX128_1

  • Opcode word: 0x100000c3
  • Primary opcode (bits 05): 4
  • Extended opcode: 195
  • Synchronising: no
Bits Field Meaning
05 OPCD primary opcode (4)
610 VD128l destination low 5 bits
1115 RA address register
1620 RB offset register
2127 XO extended opcode
2829 VD128h destination high 2 bits
3031 reserved

Operands

Field Role Description
RA0 lvx: read; lvx128: read Source GPR; when the encoded register number is 0 the operand is the literal 64-bit zero, not r0.
RB lvx: read; lvx128: read Source GPR.
VD lvx: write; lvx128: write Destination vector register.

Register Effects

lvx

  • Reads (always): RA0, RB
  • Reads (conditional): none
  • Writes (always): VD
  • Writes (conditional): none

lvx128

  • Reads (always): RA0, RB
  • Reads (conditional): none
  • Writes (always): VD
  • Writes (conditional): none

Status-Register Effects

No condition-register or status-register effects.

Operation (pseudocode)

EA <- ((RA|0) + (RB)) & ~0xF           ; align to 16
VD <- byteswap(MEM(EA, 16))

C Translation Example

/* lvx VD, RA, RB — 16-byte aligned load of a vector register      */
uint64_t base = (insn.RA == 0) ? 0 : r[insn.RA];
uint32_t ea   = (uint32_t)((base + r[insn.RB]) & ~(uint64_t)0xF);
v[insn.VD]    = mem_read_vec128_be(ea);

Implementation References

lvx

Canary emitter (frozen snapshot @ f21ebd49e9)
int InstrEmit_lvx(PPCHIRBuilder& f, const InstrData& i) {
  return InstrEmit_lvx_(f, i, i.X.RT, i.X.RA, i.X.RB);
}

// ── delegates to (src/xenia/cpu/ppc/ppc_emit_altivec.cc:133) ──
int InstrEmit_lvx_(PPCHIRBuilder& f, const InstrData& i, uint32_t vd,
                   uint32_t ra, uint32_t rb) {
  Value* ea = f.And(CalculateEA_0(f, ra, rb), f.LoadConstantInt64(~0xFull));
  f.StoreVR(vd, f.ByteSwap(f.Load(ea, VEC128_TYPE)));
  return 0;
}

lvx128

Canary emitter (frozen snapshot @ f21ebd49e9)
int InstrEmit_lvx128(PPCHIRBuilder& f, const InstrData& i) {
  return InstrEmit_lvx_(f, i, VX128_1_VD128, i.VX128_1.RA, i.VX128_1.RB);
}

// ── delegates to (src/xenia/cpu/ppc/ppc_emit_altivec.cc:133) ──
int InstrEmit_lvx_(PPCHIRBuilder& f, const InstrData& i, uint32_t vd,
                   uint32_t ra, uint32_t rb) {
  Value* ea = f.And(CalculateEA_0(f, ra, rb), f.LoadConstantInt64(~0xFull));
  f.StoreVR(vd, f.ByteSwap(f.Load(ea, VEC128_TYPE)));
  return 0;
}

Special Cases & Edge Conditions

  • Alignment is forced, not checked. The low four bits of the effective address are cleared before the load — passing an unaligned EA silently reads from EA & ~0xF rather than trapping. This differs from scalar loads (no alignment enforcement) and from lvewx etc. (which architecturally use the exact EA for lane placement).
  • Big-endian lane layout. The byte at the aligned base goes into vector lane 0 (most-significant byte); the byte at base+15 lands in lane 15. On little-endian hosts the 16-byte block is byte-swapped at the memory boundary so the PowerPC-visible layout is preserved.
  • RA0 semantics. When RA = 0, the base is the literal zero. Combined with the alignment mask this lets lvx VD, 0, RB load from RB & ~0xF.
  • No update form. Unlike scalar loads, VMX loads have no u variant that post-writes the base. Use lvxl for the cache-hint variant ("last" — the line is not expected to be reused soon).
  • VMX128 sibling (lvx128). Identical semantics; the only difference is the operand encoding. VMX128 uses a 7-bit register index split across three non-contiguous bit fields (VD128l ‖ VD128h), addressing v0..v127.
  • Atomic 16 bytes. The read is a single conceptual load; observers see either all 16 old bytes or all 16 new bytes (to the extent the surrounding cache coherency model allows).
  • Cache-line behaviour. A 16-byte aligned load fits within one Xenon 128-byte cache line; cold-line cost is one fill.
  • stvx, stvx128 — the store counterparts.
  • lvxl, lvxl128 — cache-hint "last-use" load variants.
  • lvebx, lvehx, lvewx — single-element loads at the exact (sub-aligned) address.
  • lvlx, lvrx — load-left / load-right for unaligned vector I/O (combine to read across alignment).

IBM Reference