Files
Sylpheed/tools/ppc-manual/vmx/vperm.md
sim 7f8a81b8f8
All checks were successful
CI / Native — linux (pull_request) Successful in 2h2m37s
CI / WASM — Web (pull_request) Successful in 30m13s
CI / Formatting (pull_request) Successful in 1m23s
docs(ppc-manual): quote Canary and our own decoder, not the retired xenia-rs
The generator had not been able to run correctly since the manual moved into
`tools/ppc-manual/`: it computed the repository root as `HERE.parent.parent`,
which now names `tools/`, so the XML, Canary's emitters and xenia-rs all stopped
resolving — silently, because both scrapers skipped what they could not find.
Every page's references had been pointing at paths that exist nowhere.

What each source contributed, measured on the 350 pages before this change:

  Operation (pseudocode)  251 pages: fixed boilerplate "derives from the xenia-rs
                          interpreter"; 99 carry real hand-written seeds
  C translation           337 pages: the same kind of boilerplate
  xenia-rs snapshot       336 pages: the interpreter arm, pasted in — the only
                          per-instruction semantics on unseeded pages
  links                   xenia-rs opcode/decoder/interpreter + Canary emitter

Now:

  * semantics come from **Xenia Canary**, the reference emulator, read through
    `git show` at a pinned upstream commit (`origin/canary_experimental`,
    f21ebd49e9). Not our checkout: it carries instrumentation and lacked
    upstream's `mcrf` fix, so it would have published probes and a wrong `mcrf`.
    Each page embeds the emitter (`InstrEmit_<mnem>`), and for the 128 pure
    one-line delegations also the helper that holds the semantics.
  * decode references point at `crates/sylpheed-ppc` — the decoder that
    produces `sylpheed.db` — as in-repo relative links.
  * the boilerplate now says what is true, and the C translation guide maps
    Canary's actual HIR calls, checked against `ppc_hir_builder.h` (including
    that `UpdateCR(n, v)` truncates to 32 bits).
  * `rust_scraper.py` -> `decoder_scraper.py` (interpreter half dropped);
    missing sources are now errors, not empty results.

Verified:

  consistency checks        455 XML entries, 350 families, 598 index keys
  hand-written tails        386/386 byte-identical after regeneration
  xenia-rs in generated     0
  pages with a snapshot     349/350 (was 336) — `dcbi` has no Canary emitter at all
  in-repo decoder links     910/910 resolve to a line holding the identifier
  emitter boundaries        brace counter == column-0 `}` rule on 521/521;
                            preprocessor model unit-tested (#if 0/#else/#elif)
  idempotency               re-run: 0 pages updated, 0 working-tree changes

Hand-written notes (outside the generated regions) are not rewritten here:

  * 110 links into `../../xenia-rs/...` were dead; they now point at the file in
    the archived repository (git.mc02.dev/fabi/xenia-rs @ 8401d4d). Line anchors
    were dropped because the notes predate that commit — 0 of 441 old line
    ranges match it — and a precise-looking wrong anchor is worse than none. The
    link text, which carries the author's line numbers, is unchanged.
  * 140 prose claims about xenia-rs's behaviour remain. 23 are verified to hold
    for Canary too (the 32-bit CR0 truncation, OE left unimplemented); the other
    114 need checking one by one, and some invert — e.g. `divdx` notes a correct
    64-bit CR0 update in xenia-rs where Canary's `UpdateCR` truncates. Left for
    a deliberate pass rather than a blind substitution.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 20:34:34 +02:00

8.0 KiB
Raw Blame History

vperm — Vector Permute

Category: VMX (Altivec) · Form: VA · Opcode: 0x1000002b

Assembler Mnemonics

Mnemonic XML entry Flags Description
vperm vperm — Vector Permute
vperm128 vperm128 — Vector128 Permute

Syntax

vperm [VD], [VA], [VB], [VC]
vperm128 [VD], [VA], [VB], [VC]

Encoding

vperm — form VA

  • Opcode word: 0x1000002b
  • Primary opcode (bits 0–5): 4
  • Extended opcode: 43
  • Synchronising: no
Bits Field Meaning
0–5 OPCD primary opcode (4)
6–10 VRT destination vector register
11–15 VRA source A
16–20 VRB source B
21–25 VRC source C / shift
26–31 XO extended opcode (6 bits)

vperm128 — form VX128_2

  • Opcode word: 0x14000000
  • Primary opcode (bits 0–5): 5
  • Extended opcode: 0
  • Synchronising: no
Bits Field Meaning
0–5 OPCD primary opcode (5)
6–10 VD128l destination low 5 bits
11–15 VA128l source A low 5 bits
16–20 VB128l source B low 5 bits
21 VA128H source A high bit
23–25 VC source C 3-bit field
26 VA128h source A middle bit
28–29 VD128h destination high 2 bits
30–31 VB128h source B high 2 bits

Operands

Field Role Description
VA vperm: read; vperm128: read Source A vector register.
VB vperm: read; vperm128: read Source B vector register.
VC vperm: read; vperm128: read Source C vector register / 3-bit selector.
VD vperm: write; vperm128: write Destination vector register.

Register Effects

vperm

  • Reads (always): VA, VB, VC
  • Reads (conditional): none
  • Writes (always): VD
  • Writes (conditional): none

vperm128

  • Reads (always): VA, VB, VC
  • Reads (conditional): none
  • Writes (always): VD
  • Writes (conditional): none

Status-Register Effects

No condition-register or status-register effects.

Operation (pseudocode)

; No hand-written pseudocode for this instruction yet.
; The authoritative semantics are the Canary emitter snapshot under
; Implementation References; about half of Canary's emitters open
; with the PPC-style definition as a comment (`RD <- (RA) + (RB)`).
; Every side effect is also enumerated in the Register Effects and
; Status-Register Effects tables above.

C Translation Example

/* No hand-written C yet. Translate the Canary emitter snapshot   */
/* under Implementation References; its HIR maps directly:        */
/*   f.LoadGPR(n) / f.StoreGPR(n, v)  -> r[n] / r[n] = v          */
/*   f.LoadFPR / StoreFPR, f.LoadVR / StoreVR -> f[n], v[n]        */
/*   f.Load(ea, T), f.Store(ea, v) -> raw read / write; emitters   */
/*     wrap them in f.ByteSwap for the big-endian guest value      */
/*   f.UpdateCR(n, v)  -> CR field n from v's LOW 32 BITS vs 0     */
/*   f.LoadCA / f.StoreCA -> xer.CA;  f.StoreSAT -> vscr.SAT       */
/*   i.XO.RA, i.D.DS, ... -> the bit-fields listed under Operands  */
/* The Register Effects and Status-Register Effects tables above  */
/* enumerate every side effect a faithful translation must emit.  */

Implementation References

vperm

Canary emitter (frozen snapshot @ f21ebd49e9)
int InstrEmit_vperm(PPCHIRBuilder& f, const InstrData& i) {
  return InstrEmit_vperm_(f, i.VXA.VD, i.VXA.VA, i.VXA.VB, i.VXA.VC);
}

// ── delegates to (src/xenia/cpu/ppc/ppc_emit_altivec.cc:1169) ──
int InstrEmit_vperm_(PPCHIRBuilder& f, uint32_t vd, uint32_t va, uint32_t vb,
                     uint32_t vc) {
  Value* v = f.Permute(f.LoadVR(vc), f.LoadVR(va), f.LoadVR(vb), INT8_TYPE);
  f.StoreVR(vd, v);
  return 0;
}

vperm128

Canary emitter (frozen snapshot @ f21ebd49e9)
int InstrEmit_vperm128(PPCHIRBuilder& f, const InstrData& i) {
  return InstrEmit_vperm_(f, VX128_2_VD128, VX128_2_VA128, VX128_2_VB128,
                          VX128_2_VC);
}

Special Cases & Edge Conditions

  • Per-byte selector drives a cross-vector permute. Each byte of VC is a 5-bit selector (low 5 bits used, upper 3 bits ignored). Bit 3 of that 5-bit field (i.e. the "16 bit") chooses which source: 0 selects from VA, 1 selects from VB. The low 4 bits index a byte within the chosen 16-byte operand.
  • vperm is the universal "16-byte reshuffle" primitive. It can express any byte-level permutation of 32 source bytes (VA ‖ VB) down to 16 destination bytes, including duplicates and drops.
  • Big-endian byte indexing. VC.b[0] controls VD.b[0] (the MSB byte). Selector value 0 picks VA.b[0], value 15 picks VA.b[15], value 16 picks VB.b[0], value 31 picks VB.b[15].
  • Upper 3 bits of each VC byte are ignored. Only bits 3..7 (the low 5) are consulted, so values like 0x1F and 0x5F both mean "byte 15 of VB". Software can use those upper bits for its own tagging.
  • Pair with lvsl / lvsr for unaligned 16-byte loads. lvsl produces the selector that shifts "left" by EA & 0xF bytes; feeding that into vperm with two aligned lvx results yields the unaligned 16-byte view.
  • Aliasing legal. VD may equal VA or VB.
  • VMX128 sibling vperm128. Same shape with the 7-bit register file. The VMX128 encoding carries VC in the 3-bit VC sub-field of the VX128_2 form — which only lets VC select one of 8 specific registers, not 128. In xenia's decoder this is vc128().
  • No flags, no VSCR side-effect.
  • vsldoi — static-shift-by-SHB form; when the shift is a compile-time constant this is cheaper than lvsl+vperm.
  • lvsl, lvsr — generate the permute mask from an effective address.
  • vmrghb, vmrglb, vmrghh, vmrglh, vmrghw, vmrglw — dedicated merges that are a subset of vperm.
  • vspltb, vsplth, vspltw — splat-from-lane, also expressible via vperm + a constant mask.
  • vpkuhum and other vpk* — narrower-lane packs whose pattern can also be encoded in vperm.

IBM Reference