Files
Sylpheed/tools/ppc-manual/vmx/vperm.md
MechaCat02 10dc260f0c chore(tools): adopt the PPC manual and a canary launcher that works anywhere
CONSOLIDATION.md Phase 6. Both lived untracked in the project root -- on one
disk, backed up by nothing.

  tools/ppc-manual/   393 files, 3.7 MB. 455 instructions, 350 family pages,
                      598 mnemonics resolvable through index.json, plus the
                      generator that produced them.
  tools/run-canary.sh the oracle launcher.

🔴 THE LAUNCHER WAS BROKEN IN TWO WAYS AND IS REWRITTEN, not copied:

  * it pointed at `xenia-rs/sylpheed.iso`, a SYMLINK. Wine cannot resolve one
    and says "path invalid", which reads as a corrupt image rather than a path
    problem -- it has cost a session before. It now points at the real file and
    warns if handed a symlink.
  * it hardcoded one machine's absolute paths, and named `xenia-rs`, which this
    consolidation retires. Now derived from the script's own location, with
    SYLPH_CANARY_BIN / SYLPH_ISO overrides and a check that each exists.

The standing constraints are in its header where someone will read them: one
emulator at a time, Canary runs MUTED, and never judge a crash or a hang from
a Bash-launched run -- a SIGKILL that looked like the binary was the editor's
process supervisor.

⚠️ The manual's GENERATOR reads the xenia-rs source tree, which is going away.
Its decoder now lives here as crates/sylpheed-ppc, so the generator must be
repointed before it is run again. Recorded in the README rather than left for
someone to discover; the manual's content is checked in and regenerates from
nothing implicitly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 21:18:00 +02:00

9.2 KiB
Raw Blame History

vperm — Vector Permute

Category: VMX (Altivec) · Form: VA · Opcode: 0x1000002b

Assembler Mnemonics

Mnemonic XML entry Flags Description
vperm vperm — Vector Permute
vperm128 vperm128 — Vector128 Permute

Syntax

vperm [VD], [VA], [VB], [VC]
vperm128 [VD], [VA], [VB], [VC]

Encoding

vperm — form VA

  • Opcode word: 0x1000002b
  • Primary opcode (bits 0–5): 4
  • Extended opcode: 43
  • Synchronising: no
Bits Field Meaning
0–5 OPCD primary opcode (4)
6–10 VRT destination vector register
11–15 VRA source A
16–20 VRB source B
21–25 VRC source C / shift
26–31 XO extended opcode (6 bits)

vperm128 — form VX128_2

  • Opcode word: 0x14000000
  • Primary opcode (bits 0–5): 5
  • Extended opcode: 0
  • Synchronising: no
Bits Field Meaning
0–5 OPCD primary opcode (5)
6–10 VD128l destination low 5 bits
11–15 VA128l source A low 5 bits
16–20 VB128l source B low 5 bits
21 VA128H source A high bit
23–25 VC source C 3-bit field
26 VA128h source A middle bit
28–29 VD128h destination high 2 bits
30–31 VB128h source B high 2 bits

Operands

Field Role Description
VA vperm: read; vperm128: read Source A vector register.
VB vperm: read; vperm128: read Source B vector register.
VC vperm: read; vperm128: read Source C vector register / 3-bit selector.
VD vperm: write; vperm128: write Destination vector register.

Register Effects

vperm

  • Reads (always): VA, VB, VC
  • Reads (conditional): none
  • Writes (always): VD
  • Writes (conditional): none

vperm128

  • Reads (always): VA, VB, VC
  • Reads (conditional): none
  • Writes (always): VD
  • Writes (conditional): none

Status-Register Effects

No condition-register or status-register effects.

Operation (pseudocode)

; Pseudocode derives directly from the xenia-rs interpreter
; arm (see Implementation References). Operation semantics:
;   - Read source operands from the fields listed under Operands.
;   - Apply the arithmetic / logical / memory action described
;     in the Description field above.
;   - Write results to the destination register(s); update any
;     status bits enumerated under Status-Register Effects.
; Consult the IBM AIX reference link under IBM Reference for
; canonical PPC-style pseudocode where xenia's expression is
; terse.

C Translation Example

/* C translation: the xenia-rs interpreter arm below in           */
/* Implementation References is the authoritative semantic        */
/* snapshot. Translate it line-by-line:                            */
/*   - ctx.gpr[N]  -> r[N]       (or f[]/v[] for FPRs/VRs)        */
/*   - mem.read_u*/write_u* -> mem_read_u*_be / mem_write_u*_be   */
/*   - ctx.update_cr_signed(fld, v) -> update_cr_signed(fld, v)   */
/*   - ctx.xer_ca / xer_ov / xer_so -> xer.CA / xer.OV / xer.SO   */
/* The Register Effects and Status-Register Effects tables above  */
/* enumerate every side effect a faithful translation must emit.  */

Implementation References

vperm

xenia-rs interpreter body (frozen snapshot)
        PpcOpcode::vperm | PpcOpcode::vperm128 => {
            let (va, vb, vd);
            let vc;
            if matches!(instr.opcode, PpcOpcode::vperm128) {
                va = instr.va128();
                vb = instr.vb128();
                vd = instr.vd128();
                vc = instr.vc128_2();
            } else {
                va = instr.ra();
                vb = instr.rb();
                vd = instr.rd();
                vc = instr.rc();
            }
            let a_bytes = ctx.vr[va].as_bytes();
            let b_bytes = ctx.vr[vb].as_bytes();
            let c_bytes = ctx.vr[vc].as_bytes();
            let mut r = [0u8; 16];
            for i in 0..16 {
                let idx = (c_bytes[i] & 0x1F) as usize;
                r[i] = if idx < 16 { a_bytes[idx] } else { b_bytes[idx - 16] };
            }
            ctx.vr[vd] = xenia_types::Vec128::from_bytes(r);
            ctx.pc += 4;
        }

vperm128

xenia-rs interpreter body (frozen snapshot)
        PpcOpcode::vperm | PpcOpcode::vperm128 => {
            let (va, vb, vd);
            let vc;
            if matches!(instr.opcode, PpcOpcode::vperm128) {
                va = instr.va128();
                vb = instr.vb128();
                vd = instr.vd128();
                vc = instr.vc128_2();
            } else {
                va = instr.ra();
                vb = instr.rb();
                vd = instr.rd();
                vc = instr.rc();
            }
            let a_bytes = ctx.vr[va].as_bytes();
            let b_bytes = ctx.vr[vb].as_bytes();
            let c_bytes = ctx.vr[vc].as_bytes();
            let mut r = [0u8; 16];
            for i in 0..16 {
                let idx = (c_bytes[i] & 0x1F) as usize;
                r[i] = if idx < 16 { a_bytes[idx] } else { b_bytes[idx - 16] };
            }
            ctx.vr[vd] = xenia_types::Vec128::from_bytes(r);
            ctx.pc += 4;
        }

Special Cases & Edge Conditions

  • Per-byte selector drives a cross-vector permute. Each byte of VC is a 5-bit selector (low 5 bits used, upper 3 bits ignored). Bit 3 of that 5-bit field (i.e. the "16 bit") chooses which source: 0 selects from VA, 1 selects from VB. The low 4 bits index a byte within the chosen 16-byte operand.
  • vperm is the universal "16-byte reshuffle" primitive. It can express any byte-level permutation of 32 source bytes (VA ‖ VB) down to 16 destination bytes, including duplicates and drops.
  • Big-endian byte indexing. VC.b[0] controls VD.b[0] (the MSB byte). Selector value 0 picks VA.b[0], value 15 picks VA.b[15], value 16 picks VB.b[0], value 31 picks VB.b[15].
  • Upper 3 bits of each VC byte are ignored. Only bits 3..7 (the low 5) are consulted, so values like 0x1F and 0x5F both mean "byte 15 of VB". Software can use those upper bits for its own tagging.
  • Pair with lvsl / lvsr for unaligned 16-byte loads. lvsl produces the selector that shifts "left" by EA & 0xF bytes; feeding that into vperm with two aligned lvx results yields the unaligned 16-byte view.
  • Aliasing legal. VD may equal VA or VB.
  • VMX128 sibling vperm128. Same shape with the 7-bit register file. The VMX128 encoding carries VC in the 3-bit VC sub-field of the VX128_2 form — which only lets VC select one of 8 specific registers, not 128. In xenia's decoder this is vc128().
  • No flags, no VSCR side-effect.
  • vsldoi — static-shift-by-SHB form; when the shift is a compile-time constant this is cheaper than lvsl+vperm.
  • lvsl, lvsr — generate the permute mask from an effective address.
  • vmrghb, vmrglb, vmrghh, vmrglh, vmrghw, vmrglw — dedicated merges that are a subset of vperm.
  • vspltb, vsplth, vspltw — splat-from-lane, also expressible via vperm + a constant mask.
  • vpkuhum and other vpk* — narrower-lane packs whose pattern can also be encoded in vperm.

IBM Reference