Files
Sylpheed/tools/ppc-manual/vmx/lvsr.md
sim 7f8a81b8f8
All checks were successful
CI / Native — linux (pull_request) Successful in 2h2m37s
CI / WASM — Web (pull_request) Successful in 30m13s
CI / Formatting (pull_request) Successful in 1m23s
docs(ppc-manual): quote Canary and our own decoder, not the retired xenia-rs
The generator had not been able to run correctly since the manual moved into
`tools/ppc-manual/`: it computed the repository root as `HERE.parent.parent`,
which now names `tools/`, so the XML, Canary's emitters and xenia-rs all stopped
resolving — silently, because both scrapers skipped what they could not find.
Every page's references had been pointing at paths that exist nowhere.

What each source contributed, measured on the 350 pages before this change:

  Operation (pseudocode)  251 pages: fixed boilerplate "derives from the xenia-rs
                          interpreter"; 99 carry real hand-written seeds
  C translation           337 pages: the same kind of boilerplate
  xenia-rs snapshot       336 pages: the interpreter arm, pasted in — the only
                          per-instruction semantics on unseeded pages
  links                   xenia-rs opcode/decoder/interpreter + Canary emitter

Now:

  * semantics come from **Xenia Canary**, the reference emulator, read through
    `git show` at a pinned upstream commit (`origin/canary_experimental`,
    f21ebd49e9). Not our checkout: it carries instrumentation and lacked
    upstream's `mcrf` fix, so it would have published probes and a wrong `mcrf`.
    Each page embeds the emitter (`InstrEmit_<mnem>`), and for the 128 pure
    one-line delegations also the helper that holds the semantics.
  * decode references point at `crates/sylpheed-ppc` — the decoder that
    produces `sylpheed.db` — as in-repo relative links.
  * the boilerplate now says what is true, and the C translation guide maps
    Canary's actual HIR calls, checked against `ppc_hir_builder.h` (including
    that `UpdateCR(n, v)` truncates to 32 bits).
  * `rust_scraper.py` -> `decoder_scraper.py` (interpreter half dropped);
    missing sources are now errors, not empty results.

Verified:

  consistency checks        455 XML entries, 350 families, 598 index keys
  hand-written tails        386/386 byte-identical after regeneration
  xenia-rs in generated     0
  pages with a snapshot     349/350 (was 336) — `dcbi` has no Canary emitter at all
  in-repo decoder links     910/910 resolve to a line holding the identifier
  emitter boundaries        brace counter == column-0 `}` rule on 521/521;
                            preprocessor model unit-tested (#if 0/#else/#elif)
  idempotency               re-run: 0 pages updated, 0 working-tree changes

Hand-written notes (outside the generated regions) are not rewritten here:

  * 110 links into `../../xenia-rs/...` were dead; they now point at the file in
    the archived repository (git.mc02.dev/fabi/xenia-rs @ 8401d4d). Line anchors
    were dropped because the notes predate that commit — 0 of 441 old line
    ranges match it — and a precise-looking wrong anchor is worse than none. The
    link text, which carries the author's line numbers, is unchanged.
  * 140 prose claims about xenia-rs's behaviour remain. 23 are verified to hold
    for Canary too (the 32-bit CR0 truncation, OE left unimplemented); the other
    114 need checking one by one, and some invert — e.g. `divdx` notes a correct
    64-bit CR0 update in xenia-rs where Canary's `UpdateCR` truncates. Left for
    a deliberate pass rather than a blind substitution.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 20:34:34 +02:00

8.0 KiB
Raw Blame History

lvsr — Load Vector for Shift Right Indexed

Category: VMX (Altivec) · Form: X · Opcode: 0x7c00004c

Assembler Mnemonics

Mnemonic XML entry Flags Description
lvsr lvsr — Load Vector for Shift Right Indexed
lvsr128 lvsr128 — Load Vector for Shift Right Indexed 128

Syntax

lvsr [VD], [RA0], [RB]
lvsr128 [VD], [RA0], [RB]

Encoding

lvsr — form X

  • Opcode word: 0x7c00004c
  • Primary opcode (bits 0–5): 31
  • Extended opcode: 38
  • Synchronising: no
Bits Field Meaning
0–5 OPCD primary opcode
6–10 RT/FRT/VRT destination
11–15 RA/FRA/VRA source A
16–20 RB/FRB/VRB source B
21–30 XO extended opcode (10 bits)
31 Rc record-form flag

lvsr128 — form VX128_1

  • Opcode word: 0x10000043
  • Primary opcode (bits 0–5): 4
  • Extended opcode: 67
  • Synchronising: no
Bits Field Meaning
0–5 OPCD primary opcode (4)
6–10 VD128l destination low 5 bits
11–15 RA address register
16–20 RB offset register
21–27 XO extended opcode
28–29 VD128h destination high 2 bits
30–31 — reserved

Operands

Field Role Description
RA0 lvsr: read; lvsr128: read Source GPR; when the encoded register number is 0 the operand is the literal 64-bit zero, not r0.
RB lvsr: read; lvsr128: read Source GPR.
VD lvsr: write; lvsr128: write Destination vector register.

Register Effects

lvsr

  • Reads (always): RA0, RB
  • Reads (conditional): none
  • Writes (always): VD
  • Writes (conditional): none

lvsr128

  • Reads (always): RA0, RB
  • Reads (conditional): none
  • Writes (always): VD
  • Writes (conditional): none

Status-Register Effects

No condition-register or status-register effects.

Operation (pseudocode)

addr_lo <- ((RA|0) + (RB))[60:63]
for i in 0..15: VD[i] <- 16 − addr_lo + i

C Translation Example

/* No hand-written C yet. Translate the Canary emitter snapshot   */
/* under Implementation References; its HIR maps directly:        */
/*   f.LoadGPR(n) / f.StoreGPR(n, v)  -> r[n] / r[n] = v          */
/*   f.LoadFPR / StoreFPR, f.LoadVR / StoreVR -> f[n], v[n]        */
/*   f.Load(ea, T), f.Store(ea, v) -> raw read / write; emitters   */
/*     wrap them in f.ByteSwap for the big-endian guest value      */
/*   f.UpdateCR(n, v)  -> CR field n from v's LOW 32 BITS vs 0     */
/*   f.LoadCA / f.StoreCA -> xer.CA;  f.StoreSAT -> vscr.SAT       */
/*   i.XO.RA, i.D.DS, ... -> the bit-fields listed under Operands  */
/* The Register Effects and Status-Register Effects tables above  */
/* enumerate every side effect a faithful translation must emit.  */

Implementation References

lvsr

Canary emitter (frozen snapshot @ f21ebd49e9)
int InstrEmit_lvsr(PPCHIRBuilder& f, const InstrData& i) {
  return InstrEmit_lvsr_(f, i, i.X.RT, i.X.RA, i.X.RB);
}

// ── delegates to (src/xenia/cpu/ppc/ppc_emit_altivec.cc:118) ──
int InstrEmit_lvsr_(PPCHIRBuilder& f, const InstrData& i, uint32_t vd,
                    uint32_t ra, uint32_t rb) {
  Value* ea = CalculateEA_0(f, ra, rb);
  Value* sh = f.Truncate(f.And(ea, f.LoadConstantInt64(0xF)), INT8_TYPE);
  Value* v = f.LoadVectorShr(sh);
  f.StoreVR(vd, v);
  return 0;
}

lvsr128

Canary emitter (frozen snapshot @ f21ebd49e9)
int InstrEmit_lvsr128(PPCHIRBuilder& f, const InstrData& i) {
  return InstrEmit_lvsr_(f, i, VX128_1_VD128, i.VX128_1.RA, i.VX128_1.RB);
}

// ── delegates to (src/xenia/cpu/ppc/ppc_emit_altivec.cc:118) ──
int InstrEmit_lvsr_(PPCHIRBuilder& f, const InstrData& i, uint32_t vd,
                    uint32_t ra, uint32_t rb) {
  Value* ea = CalculateEA_0(f, ra, rb);
  Value* sh = f.Truncate(f.And(ea, f.LoadConstantInt64(0xF)), INT8_TYPE);
  Value* v = f.LoadVectorShr(sh);
  f.StoreVR(vd, v);
  return 0;
}

Special Cases & Edge Conditions

  • No memory access. Like lvsl, lvsr does not touch memory: the effective address is consumed solely to extract the low four bits, which then drive the synthesised permute mask in VD.
  • Mirror of lvsl. Where lvsl produces {sh, sh+1, …, sh+15}, lvsr produces {16−sh, 17−sh, …, 31−sh}. When EA & 0xF == 0 the output is {16, 17, …, 31} — the identity permute that selects all of VB (in the vperm VD, VA, VB, VC orientation). When EA & 0xF == 3 the output is {13, 14, …, 28}, splitting the vperm between the high three bytes of VA and the low thirteen of VB.
  • Big-endian byte indexing. VD[0] is the most-significant byte (the byte at the lowest address after a stvx).
  • Right-shift unaligned-load idiom. Pair with two aligned lvx and a vperm when the source data is laid out so the wanted vector starts in the second aligned block:
    lvx   vAL, r0, rA           ; aligned block at EA & ~0xF
    lvx   vAH, r0, rA + 16      ; next aligned block
    lvsr  vC,  r0, rA           ; right-shift permute mask
    vperm vD,  vAH, vAL, vC     ; note: vAH then vAL — opposite of lvsl
    
    The argument flip versus the lvsl idiom is the whole reason both masks exist.
  • RA0 semantics. When RA = 0 the base is the literal zero, so lvsr vD, 0, rB derives the mask from rB & 0xF.
  • Selectors >15 are intentional. Inside vperm, byte selectors with bit 4 set (i.e. >= 16) index into the second source vector. lvsr deliberately produces values up to 31, since only the low five bits are honoured by vperm.
  • VMX128 sibling (lvsr128). Identical semantics; the extended VD128l ‖ VD128h encoding lets vD reach v0..v127.
  • No flags, no exceptions, trivially reorderable.
  • lvsl — the mirror: VD[i] = sh + i.
  • vperm — consumes the mask to perform arbitrary byte-level permutation across two vectors.
  • lvx, lvlx, lvrx — the actual memory loads that supply the two aligned halves.
  • vsldoi — when the misalignment is a compile-time constant, the static-offset shift is cheaper than the lvsr/vperm pair.

IBM Reference