The hand-written parts of the manual still described how the retired xenia-rs interpreter behaved: its snapshots, Rust casts and helpers. Each of those 490 statements is now either restated as what Canary's emitters and x64 backend actually do (at the pinned canary_experimental commit), or dropped where it only made sense for xenia-rs. Checking them turned up claims that were wrong, not just outdated: - VSCR[SAT] is never modelled in Canary (DID_SATURATE is a stub and mfvscr cannot see it); the pages said saturating ops set it stickily. - Canary does not implement lswi/lswx/stswi/stswx, dcbi, mtfsb0/mtfsb1, vmsum*, vmhaddshs, vupkhpx/vupklpx, and most SPRs; pages described them as working. - Traps evaluate TO in Canary; stvebx/stvehx/stvewx store one element, not 16 bytes; mtmsrd writes only EE; fres/frsqrte/vrsqrtefp precision claims and the stfs "rounds under RN / sets FPSCR" claim contradicted the spec. - Reservations are a 64 KiB block bitmap plus a value compare, not per-address tracking. Claims that neither Canary's source nor a public spec settles are marked unverified (NI at boot, vmaddcfp128 operand order, estimate bit-exactness). Generated regions are untouched; re-running the generator changes nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
6.8 KiB
6.8 KiB
vrlimi128 — Vector128 Rotate Left Immediate and Mask Insert
Assembler Mnemonics
| Mnemonic | XML entry | Flags | Description |
|---|---|---|---|
vrlimi128 |
vrlimi128 |
— | Vector128 Rotate Left Immediate and Mask Insert |
Syntax
vrlimi128 [VD], [VB], [IMM], [z]
Encoding
vrlimi128 — form VX128_4
- Opcode word:
0x18000710 - Primary opcode (bits 0–5):
6 - Extended opcode:
1808 - Synchronising: no
| Bits | Field | Meaning |
|---|---|---|
| 0–5 | OPCD |
primary opcode (6) |
| 6–10 | VD128l |
destination low 5 bits |
| 11–15 | IMM |
5-bit immediate |
| 16–20 | VB128l |
source B low 5 bits |
| 21–23 | XO |
extended opcode |
| 24–25 | z |
sub-operation selector |
| 28–29 | VD128h |
destination high 2 bits |
| 30–31 | VB128h |
source B high 2 bits |
Operands
| Field | Role | Description |
|---|---|---|
VB |
vrlimi128: read | Source B vector register. |
VD |
vrlimi128: write | Destination vector register. |
Register Effects
vrlimi128
- Reads (always):
VB - Reads (conditional): none
- Writes (always):
VD - Writes (conditional): none
Status-Register Effects
No condition-register or status-register effects.
Operation (pseudocode)
; No hand-written pseudocode for this instruction yet.
; The authoritative semantics are the Canary emitter snapshot under
; Implementation References; about half of Canary's emitters open
; with the PPC-style definition as a comment (`RD <- (RA) + (RB)`).
; Every side effect is also enumerated in the Register Effects and
; Status-Register Effects tables above.
C Translation Example
/* No hand-written C yet. Translate the Canary emitter snapshot */
/* under Implementation References; its HIR maps directly: */
/* f.LoadGPR(n) / f.StoreGPR(n, v) -> r[n] / r[n] = v */
/* f.LoadFPR / StoreFPR, f.LoadVR / StoreVR -> f[n], v[n] */
/* f.Load(ea, T), f.Store(ea, v) -> raw read / write; emitters */
/* wrap them in f.ByteSwap for the big-endian guest value */
/* f.UpdateCR(n, v) -> CR field n from v's LOW 32 BITS vs 0 */
/* f.LoadCA / f.StoreCA -> xer.CA; f.StoreSAT -> vscr.SAT */
/* i.XO.RA, i.D.DS, ... -> the bit-fields listed under Operands */
/* The Register Effects and Status-Register Effects tables above */
/* enumerate every side effect a faithful translation must emit. */
Implementation References
vrlimi128
- Canary XML:
tools/ppc-instructions.xml— search formnem="vrlimi128" - Canary emitter:
src/xenia/cpu/ppc/ppc_emit_altivec.cc:1291 - Sylpheed opcode:
crates/sylpheed-ppc/src/opcode.rs:433 - Sylpheed decoder:
crates/sylpheed-ppc/src/decoder.rs:764
Canary emitter (frozen snapshot @ f21ebd49e9)
int InstrEmit_vrlimi128(PPCHIRBuilder& f, const InstrData& i) {
const uint32_t vd = i.VX128_4.VD128l | (i.VX128_4.VD128h << 5);
const uint32_t vb = i.VX128_4.VB128l | (i.VX128_4.VB128h << 5);
uint32_t blend_mask_src = i.VX128_4.IMM;
uint32_t blend_mask = 0;
blend_mask |= (((blend_mask_src >> 3) & 0x1) ? 0 : 4) << 0;
blend_mask |= (((blend_mask_src >> 2) & 0x1) ? 1 : 5) << 8;
blend_mask |= (((blend_mask_src >> 1) & 0x1) ? 2 : 6) << 16;
blend_mask |= (((blend_mask_src >> 0) & 0x1) ? 3 : 7) << 24;
uint32_t rotate = i.VX128_4.z;
// This is just a fancy permute.
// X Y Z W, rotated left by 2 = Z W X Y
// Then mask select the results into the dest.
// Sometimes rotation is zero, so fast path.
Value* v;
if (rotate) {
// TODO(benvanik): constants need conversion.
uint32_t swizzle_mask;
switch (rotate) {
case 1:
// X Y Z W -> Y Z W X
swizzle_mask = SWIZZLE_XYZW_TO_YZWX;
break;
case 2:
// X Y Z W -> Z W X Y
swizzle_mask = SWIZZLE_XYZW_TO_ZWXY;
break;
case 3:
// X Y Z W -> W X Y Z
swizzle_mask = SWIZZLE_XYZW_TO_WXYZ;
break;
default:
XEINSTRNOTIMPLEMENTED();
return 1;
}
v = f.Swizzle(f.LoadVR(vb), FLOAT32_TYPE, swizzle_mask);
} else {
v = f.LoadVR(vb);
}
if (blend_mask != kIdentityPermuteMask) {
v = f.Permute(f.LoadConstantUint32(blend_mask), v, f.LoadVR(vd),
INT32_TYPE);
}
f.StoreVR(vd, v);
return 0;
}
Special Cases & Edge Conditions
- Rotate-left-word + mask-insert in one step.
VBis rotated left byzword positions (word-granular, 0..3 — not bits). The resulting rotated vector is merged into the pre-existingVDunder control of a 4-bit "insert mask" (Canary'sVX128_4.IMM, most-significant bit for lane 0): mask bit = 1 takes the lane from the rotatedVB; mask bit = 0 keeps the lane from the oldVD. - Destructive destination.
VDis both source and destination — software must preserve its value or pre-initialise it. - Typical use: selective-lane overwrite. Games use this to "rewrite lane
nof a vector with a shuffled component" without a full permute. A common pattern is "insert a scalar into laneiof a vector" where the scalar has been pre-loaded to a known word ofVB. - Mask bit ↔ lane mapping. Big-endian: mask bit 3 (MSB of the 4-bit mask) controls lane 0; bit 0 controls lane 3. (In Canary: lane
itakes the rotatedVBwhen(IMM >> (3 − i)) & 1.) - VMX128 register-fusion on
VDandVB. - No IBM AIX entry — Xenon-only.
- No
Rc, no XER, no VSCR.
Related Instructions
vrlw,vrlw128— per-lane bit-level rotate (word-granular shift, not lane-granular).vpermwi128— immediate 4-way word permute (no merge).vsel,vsel128— general bit-select;vrlimi128is the specialised "rotate + insert" equivalent.vsldoi— byte-level immediate shift.
IBM Reference
- No IBM AIX entry — this instruction is exclusive to the Xbox 360's VMX128 extension. The mnemonic is an adaptation of the scalar
rlwimi(rotate-left-word-immediate-mask-insert) pattern for vectors. - Xbox 360 XDK, Altivec-128 (VMX128) extensions.
- IBM AltiVec Technology Programmer's Interface Manual, Chapter 4 — Integer Shift / Rotate for the base rotate semantics.