Files
Sylpheed/tools/ppc-manual/vmx/vperm.md
sim f3c512f2ab docs(ppc-manual): check every xenia-rs claim against Canary's source
The hand-written parts of the manual still described how the retired
xenia-rs interpreter behaved: its snapshots, Rust casts and helpers. Each of
those 490 statements is now either restated as what Canary's emitters and
x64 backend actually do (at the pinned canary_experimental commit), or
dropped where it only made sense for xenia-rs.

Checking them turned up claims that were wrong, not just outdated:

- VSCR[SAT] is never modelled in Canary (DID_SATURATE is a stub and mfvscr
  cannot see it); the pages said saturating ops set it stickily.
- Canary does not implement lswi/lswx/stswi/stswx, dcbi, mtfsb0/mtfsb1,
  vmsum*, vmhaddshs, vupkhpx/vupklpx, and most SPRs; pages described them
  as working.
- Traps evaluate TO in Canary; stvebx/stvehx/stvewx store one element, not
  16 bytes; mtmsrd writes only EE; fres/frsqrte/vrsqrtefp precision claims
  and the stfs "rounds under RN / sets FPSCR" claim contradicted the spec.
- Reservations are a 64 KiB block bitmap plus a value compare, not
  per-address tracking.

Claims that neither Canary's source nor a public spec settles are marked
unverified (NI at boot, vmaddcfp128 operand order, estimate bit-exactness).

Generated regions are untouched; re-running the generator changes nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 21:52:38 +02:00

8.0 KiB
Raw Blame History

vperm — Vector Permute

Category: VMX (Altivec) · Form: VA · Opcode: 0x1000002b

Assembler Mnemonics

Mnemonic XML entry Flags Description
vperm vperm Vector Permute
vperm128 vperm128 Vector128 Permute

Syntax

vperm [VD], [VA], [VB], [VC]
vperm128 [VD], [VA], [VB], [VC]

Encoding

vperm — form VA

  • Opcode word: 0x1000002b
  • Primary opcode (bits 05): 4
  • Extended opcode: 43
  • Synchronising: no
Bits Field Meaning
05 OPCD primary opcode (4)
610 VRT destination vector register
1115 VRA source A
1620 VRB source B
2125 VRC source C / shift
2631 XO extended opcode (6 bits)

vperm128 — form VX128_2

  • Opcode word: 0x14000000
  • Primary opcode (bits 05): 5
  • Extended opcode: 0
  • Synchronising: no
Bits Field Meaning
05 OPCD primary opcode (5)
610 VD128l destination low 5 bits
1115 VA128l source A low 5 bits
1620 VB128l source B low 5 bits
21 VA128H source A high bit
2325 VC source C 3-bit field
26 VA128h source A middle bit
2829 VD128h destination high 2 bits
3031 VB128h source B high 2 bits

Operands

Field Role Description
VA vperm: read; vperm128: read Source A vector register.
VB vperm: read; vperm128: read Source B vector register.
VC vperm: read; vperm128: read Source C vector register / 3-bit selector.
VD vperm: write; vperm128: write Destination vector register.

Register Effects

vperm

  • Reads (always): VA, VB, VC
  • Reads (conditional): none
  • Writes (always): VD
  • Writes (conditional): none

vperm128

  • Reads (always): VA, VB, VC
  • Reads (conditional): none
  • Writes (always): VD
  • Writes (conditional): none

Status-Register Effects

No condition-register or status-register effects.

Operation (pseudocode)

; No hand-written pseudocode for this instruction yet.
; The authoritative semantics are the Canary emitter snapshot under
; Implementation References; about half of Canary's emitters open
; with the PPC-style definition as a comment (`RD <- (RA) + (RB)`).
; Every side effect is also enumerated in the Register Effects and
; Status-Register Effects tables above.

C Translation Example

/* No hand-written C yet. Translate the Canary emitter snapshot   */
/* under Implementation References; its HIR maps directly:        */
/*   f.LoadGPR(n) / f.StoreGPR(n, v)  -> r[n] / r[n] = v          */
/*   f.LoadFPR / StoreFPR, f.LoadVR / StoreVR -> f[n], v[n]        */
/*   f.Load(ea, T), f.Store(ea, v) -> raw read / write; emitters   */
/*     wrap them in f.ByteSwap for the big-endian guest value      */
/*   f.UpdateCR(n, v)  -> CR field n from v's LOW 32 BITS vs 0     */
/*   f.LoadCA / f.StoreCA -> xer.CA;  f.StoreSAT -> vscr.SAT       */
/*   i.XO.RA, i.D.DS, ... -> the bit-fields listed under Operands  */
/* The Register Effects and Status-Register Effects tables above  */
/* enumerate every side effect a faithful translation must emit.  */

Implementation References

vperm

Canary emitter (frozen snapshot @ f21ebd49e9)
int InstrEmit_vperm(PPCHIRBuilder& f, const InstrData& i) {
  return InstrEmit_vperm_(f, i.VXA.VD, i.VXA.VA, i.VXA.VB, i.VXA.VC);
}

// ── delegates to (src/xenia/cpu/ppc/ppc_emit_altivec.cc:1169) ──
int InstrEmit_vperm_(PPCHIRBuilder& f, uint32_t vd, uint32_t va, uint32_t vb,
                     uint32_t vc) {
  Value* v = f.Permute(f.LoadVR(vc), f.LoadVR(va), f.LoadVR(vb), INT8_TYPE);
  f.StoreVR(vd, v);
  return 0;
}

vperm128

Canary emitter (frozen snapshot @ f21ebd49e9)
int InstrEmit_vperm128(PPCHIRBuilder& f, const InstrData& i) {
  return InstrEmit_vperm_(f, VX128_2_VD128, VX128_2_VA128, VX128_2_VB128,
                          VX128_2_VC);
}

Special Cases & Edge Conditions

  • Per-byte selector drives a cross-vector permute. Each byte of VC is a 5-bit selector (low 5 bits used, upper 3 bits ignored). Bit 3 of that 5-bit field (i.e. the "16 bit") chooses which source: 0 selects from VA, 1 selects from VB. The low 4 bits index a byte within the chosen 16-byte operand.
  • vperm is the universal "16-byte reshuffle" primitive. It can express any byte-level permutation of 32 source bytes (VA ‖ VB) down to 16 destination bytes, including duplicates and drops.
  • Big-endian byte indexing. VC.b[0] controls VD.b[0] (the MSB byte). Selector value 0 picks VA.b[0], value 15 picks VA.b[15], value 16 picks VB.b[0], value 31 picks VB.b[15].
  • Upper 3 bits of each VC byte are ignored. Only bits 3..7 (the low 5) are consulted, so values like 0x1F and 0x5F both mean "byte 15 of VB". Software can use those upper bits for its own tagging.
  • Pair with lvsl / lvsr for unaligned 16-byte loads. lvsl produces the selector that shifts "left" by EA & 0xF bytes; feeding that into vperm with two aligned lvx results yields the unaligned 16-byte view.
  • Aliasing legal. VD may equal VA or VB.
  • VMX128 sibling vperm128. Same shape with the 7-bit register file. The VMX128 encoding carries VC in the 3-bit VC sub-field of the VX128_2 form — which only lets VC select one of 8 specific registers, not 128. Canary reads it as VX128_2.VC.
  • No flags, no VSCR side-effect.
  • vsldoi — static-shift-by-SHB form; when the shift is a compile-time constant this is cheaper than lvsl+vperm.
  • lvsl, lvsr — generate the permute mask from an effective address.
  • vmrghb, vmrglb, vmrghh, vmrglh, vmrghw, vmrglw — dedicated merges that are a subset of vperm.
  • vspltb, vsplth, vspltw — splat-from-lane, also expressible via vperm + a constant mask.
  • vpkuhum and other vpk* — narrower-lane packs whose pattern can also be encoded in vperm.

IBM Reference