The hand-written parts of the manual still described how the retired xenia-rs interpreter behaved: its snapshots, Rust casts and helpers. Each of those 490 statements is now either restated as what Canary's emitters and x64 backend actually do (at the pinned canary_experimental commit), or dropped where it only made sense for xenia-rs. Checking them turned up claims that were wrong, not just outdated: - VSCR[SAT] is never modelled in Canary (DID_SATURATE is a stub and mfvscr cannot see it); the pages said saturating ops set it stickily. - Canary does not implement lswi/lswx/stswi/stswx, dcbi, mtfsb0/mtfsb1, vmsum*, vmhaddshs, vupkhpx/vupklpx, and most SPRs; pages described them as working. - Traps evaluate TO in Canary; stvebx/stvehx/stvewx store one element, not 16 bytes; mtmsrd writes only EE; fres/frsqrte/vrsqrtefp precision claims and the stfs "rounds under RN / sets FPSCR" claim contradicted the spec. - Reservations are a 64 KiB block bitmap plus a value compare, not per-address tracking. Claims that neither Canary's source nor a public spec settles are marked unverified (NI at boot, vmaddcfp128 operand order, estimate bit-exactness). Generated regions are untouched; re-running the generator changes nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
7.6 KiB
7.6 KiB
lvewx — Load Vector Element Word Indexed
Assembler Mnemonics
| Mnemonic | XML entry | Flags | Description |
|---|---|---|---|
lvewx |
lvewx |
— | Load Vector Element Word Indexed |
lvewx128 |
lvewx128 |
— | Load Vector Element Word Indexed 128 |
Syntax
lvewx [VD], [RA0], [RB]
lvewx128 [VD], [RA0], [RB]
Encoding
lvewx — form X
- Opcode word:
0x7c00008e - Primary opcode (bits 0–5):
31 - Extended opcode:
71 - Synchronising: no
| Bits | Field | Meaning |
|---|---|---|
| 0–5 | OPCD |
primary opcode |
| 6–10 | RT/FRT/VRT |
destination |
| 11–15 | RA/FRA/VRA |
source A |
| 16–20 | RB/FRB/VRB |
source B |
| 21–30 | XO |
extended opcode (10 bits) |
| 31 | Rc |
record-form flag |
lvewx128 — form VX128_1
- Opcode word:
0x10000083 - Primary opcode (bits 0–5):
4 - Extended opcode:
131 - Synchronising: no
| Bits | Field | Meaning |
|---|---|---|
| 0–5 | OPCD |
primary opcode (4) |
| 6–10 | VD128l |
destination low 5 bits |
| 11–15 | RA |
address register |
| 16–20 | RB |
offset register |
| 21–27 | XO |
extended opcode |
| 28–29 | VD128h |
destination high 2 bits |
| 30–31 | — |
reserved |
Operands
| Field | Role | Description |
|---|---|---|
RA0 |
lvewx: read; lvewx128: read | Source GPR; when the encoded register number is 0 the operand is the literal 64-bit zero, not r0. |
RB |
lvewx: read; lvewx128: read | Source GPR. |
VD |
lvewx: write; lvewx128: write | Destination vector register. |
Register Effects
lvewx
- Reads (always):
RA0,RB - Reads (conditional): none
- Writes (always):
VD - Writes (conditional): none
lvewx128
- Reads (always):
RA0,RB - Reads (conditional): none
- Writes (always):
VD - Writes (conditional): none
Status-Register Effects
No condition-register or status-register effects.
Operation (pseudocode)
; No hand-written pseudocode for this instruction yet.
; The authoritative semantics are the Canary emitter snapshot under
; Implementation References; about half of Canary's emitters open
; with the PPC-style definition as a comment (`RD <- (RA) + (RB)`).
; Every side effect is also enumerated in the Register Effects and
; Status-Register Effects tables above.
C Translation Example
/* No hand-written C yet. Translate the Canary emitter snapshot */
/* under Implementation References; its HIR maps directly: */
/* f.LoadGPR(n) / f.StoreGPR(n, v) -> r[n] / r[n] = v */
/* f.LoadFPR / StoreFPR, f.LoadVR / StoreVR -> f[n], v[n] */
/* f.Load(ea, T), f.Store(ea, v) -> raw read / write; emitters */
/* wrap them in f.ByteSwap for the big-endian guest value */
/* f.UpdateCR(n, v) -> CR field n from v's LOW 32 BITS vs 0 */
/* f.LoadCA / f.StoreCA -> xer.CA; f.StoreSAT -> vscr.SAT */
/* i.XO.RA, i.D.DS, ... -> the bit-fields listed under Operands */
/* The Register Effects and Status-Register Effects tables above */
/* enumerate every side effect a faithful translation must emit. */
Implementation References
lvewx
- Canary XML:
tools/ppc-instructions.xml— search formnem="lvewx" - Canary emitter:
src/xenia/cpu/ppc/ppc_emit_altivec.cc:96 - Sylpheed opcode:
crates/sylpheed-ppc/src/opcode.rs:138 - Sylpheed decoder:
crates/sylpheed-ppc/src/decoder.rs:885
Canary emitter (frozen snapshot @ f21ebd49e9)
int InstrEmit_lvewx(PPCHIRBuilder& f, const InstrData& i) {
return InstrEmit_lvewx_(f, i, i.X.RT, i.X.RA, i.X.RB);
}
// ── delegates to (src/xenia/cpu/ppc/ppc_emit_altivec.cc:89) ──
int InstrEmit_lvewx_(PPCHIRBuilder& f, const InstrData& i, uint32_t vd,
uint32_t ra, uint32_t rb) {
// Same as lvx.
Value* ea = f.And(CalculateEA_0(f, ra, rb), f.LoadConstantUint64(~0xFull));
f.StoreVR(vd, f.ByteSwap(f.Load(ea, VEC128_TYPE)));
return 0;
}
lvewx128
- Canary XML:
tools/ppc-instructions.xml— search formnem="lvewx128" - Canary emitter:
src/xenia/cpu/ppc/ppc_emit_altivec.cc:99 - Sylpheed opcode:
crates/sylpheed-ppc/src/opcode.rs:139 - Sylpheed decoder:
crates/sylpheed-ppc/src/decoder.rs:529
Canary emitter (frozen snapshot @ f21ebd49e9)
int InstrEmit_lvewx128(PPCHIRBuilder& f, const InstrData& i) {
return InstrEmit_lvewx_(f, i, VX128_1_VD128, i.VX128_1.RA, i.VX128_1.RB);
}
// ── delegates to (src/xenia/cpu/ppc/ppc_emit_altivec.cc:89) ──
int InstrEmit_lvewx_(PPCHIRBuilder& f, const InstrData& i, uint32_t vd,
uint32_t ra, uint32_t rb) {
// Same as lvx.
Value* ea = f.And(CalculateEA_0(f, ra, rb), f.LoadConstantUint64(~0xFull));
f.StoreVR(vd, f.ByteSwap(f.Load(ea, VEC128_TYPE)));
return 0;
}
Special Cases & Edge Conditions
- Single word element load. Architecturally
lvewxloads exactly four bytes fromEA(which must be 4-byte aligned) and places them in the word lane(EA mod 16) >> 2of the destination vector; the other 3 word lanes are undefined. - EA must be word-aligned. The low two bits of
EAare masked by hardware. ⚠️ Canary treatslvewxandlvewx128exactly likelvx: it masks to 16-byte alignment and loads the whole 128-bit vector. - Canary simplification — full-line read. Canary's
lvewxandlvewx128both load the full aligned 16 bytes fromea & ~0xFinto the destination vector. Architectural undefined lanes are filled deterministically. RA0semantics. WhenRA = 0, base is literal zero.- No update form. No
lvewuxexists. - VMX128 sibling.
lvewx128shares semantics; the only difference is the operand encoding. VMX128 uses a 7-bit register index split acrossVD128l ‖ VD128hso it can addressv0..v127instead of the 32-register Altivec space. - Big-endian word within the lane. The byte at the lower address is the most-significant byte of the word lane.
- Common idiom. Pair with
vspltwto broadcast the loaded word to all four lanes, or withvpermto gather words from sparse memory into one vector.
Related Instructions
lvebx,lvehx— byte and half element loads.lvx,lvxl— full 16-byte aligned vector loads.lvlx,lvrx— load-left / load-right partial-vector ops.stvewx— symmetric single-word store.
IBM Reference
- AIX 7.3 —
lvewx(Load Vector Element Word Indexed) PowerISA v2.07B Book I"Vector Facility" § "Vector Load and Store" for lane-placement rules; Microsoft XDK forlvewx128.