Files
Sylpheed/tools/ppc-manual/alu/mullwx.md
sim f3c512f2ab docs(ppc-manual): check every xenia-rs claim against Canary's source
The hand-written parts of the manual still described how the retired
xenia-rs interpreter behaved: its snapshots, Rust casts and helpers. Each of
those 490 statements is now either restated as what Canary's emitters and
x64 backend actually do (at the pinned canary_experimental commit), or
dropped where it only made sense for xenia-rs.

Checking them turned up claims that were wrong, not just outdated:

- VSCR[SAT] is never modelled in Canary (DID_SATURATE is a stub and mfvscr
  cannot see it); the pages said saturating ops set it stickily.
- Canary does not implement lswi/lswx/stswi/stswx, dcbi, mtfsb0/mtfsb1,
  vmsum*, vmhaddshs, vupkhpx/vupklpx, and most SPRs; pages described them
  as working.
- Traps evaluate TO in Canary; stvebx/stvehx/stvewx store one element, not
  16 bytes; mtmsrd writes only EE; fres/frsqrte/vrsqrtefp precision claims
  and the stfs "rounds under RN / sets FPSCR" claim contradicted the spec.
- Reservations are a 64 KiB block bitmap plus a value compare, not
  per-address tracking.

Claims that neither Canary's source nor a public spec settles are marked
unverified (NI at boot, vmaddcfp128 operand order, estimate bit-exactness).

Generated regions are untouched; re-running the generator changes nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 21:52:38 +02:00

6.5 KiB
Raw Permalink Blame History

mullwx — Multiply Low Word

Category: Integer ALU · Form: XO · Opcode: 0x7c0001d6

Assembler Mnemonics

Mnemonic XML entry Flags Description
mullw mullwx — Multiply Low Word
mullwo mullwx OE=1 Multiply Low Word
mullw. mullwx Rc=1 Multiply Low Word
mullwo. mullwx OE=1, Rc=1 Multiply Low Word

Syntax

mullw[OE][Rc] [RD], [RA], [RB]

Encoding

mullwx — form XO

  • Opcode word: 0x7c0001d6
  • Primary opcode (bits 0–5): 31
  • Extended opcode: 235
  • Synchronising: no
Bits Field Meaning
0–5 OPCD primary opcode (31)
6–10 RT destination GPR
11–15 RA source A
16–20 RB source B
21 OE overflow-enable flag
22–30 XO extended opcode (9 bits)
31 Rc record-form flag

Operands

Field Role Description
RA mullwx: read Source GPR (r0–r31).
RB mullwx: read Source GPR.
RD mullwx: write Destination GPR.
CR mullwx: write (conditional) Condition-register update. When Rc=1, CR field 0 (or CR6 for vector compares, CR1 for FPU) is updated from the result.
OE mullwx: write (conditional) Overflow-enable bit. When 1, the instruction updates XER[OV] and stickies XER[SO] on signed overflow.

Register Effects

mullwx

  • Reads (always): RA, RB
  • Reads (conditional): none
  • Writes (always): RD
  • Writes (conditional): CR, OE

Status-Register Effects

  • mullwx: CR0 ← signed-compare(result, 0) with SO ← XER[SO], when Rc=1.; XER[OV] ← signed-overflow(result); XER[SO] stickies, when OE=1.

Operation (pseudocode)

RT <- ((RA)[32:63]) * ((RB)[32:63])    ; signed 32×32 → 64

C Translation Example

/* No hand-written C yet. Translate the Canary emitter snapshot   */
/* under Implementation References; its HIR maps directly:        */
/*   f.LoadGPR(n) / f.StoreGPR(n, v)  -> r[n] / r[n] = v          */
/*   f.LoadFPR / StoreFPR, f.LoadVR / StoreVR -> f[n], v[n]        */
/*   f.Load(ea, T), f.Store(ea, v) -> raw read / write; emitters   */
/*     wrap them in f.ByteSwap for the big-endian guest value      */
/*   f.UpdateCR(n, v)  -> CR field n from v's LOW 32 BITS vs 0     */
/*   f.LoadCA / f.StoreCA -> xer.CA;  f.StoreSAT -> vscr.SAT       */
/*   i.XO.RA, i.D.DS, ... -> the bit-fields listed under Operands  */
/* The Register Effects and Status-Register Effects tables above  */
/* enumerate every side effect a faithful translation must emit.  */

Implementation References

mullwx

Canary emitter (frozen snapshot @ f21ebd49e9)
int InstrEmit_mullwx(PPCHIRBuilder& f, const InstrData& i) {
  // RT <- (RA)[32:63] × (RB)[32:63]
  if (i.XO.OE) {
    // With XER update.
    XEINSTRNOTIMPLEMENTED();
  }
  Value* v = f.Mul(
      f.SignExtend(f.Truncate(f.LoadGPR(i.XO.RA), INT32_TYPE), INT64_TYPE),
      f.SignExtend(f.Truncate(f.LoadGPR(i.XO.RB), INT32_TYPE), INT64_TYPE));
  f.StoreGPR(i.XO.RT, v);
  if (i.XO.Rc) {
    f.UpdateCR(0, v);
  }
  return 0;
}

Extended Pseudocode

prod64 <- sign_extend_32_to_64((RA)[32:63]) *s sign_extend_32_to_64((RB)[32:63])
RT <- prod64                                         ; 64-bit result
if OE then
    XER[OV] <- (prod64 ≠ sign_extend_32_to_64(prod64[32:63]))  ; set when product doesn't fit in 32 bits
    XER[SO] <- XER[SO] | XER[OV]
if Rc then
    CR0 <- signed_compare(RT, 0) || XER[SO]

Special Cases & Edge Conditions

  • Inputs are the low 32 bits. mullw only looks at RA[32:63] and RB[32:63]; the high 32 bits of each source are ignored. This is a 32-bit × 32-bit → 64-bit signed multiply. For full 64-bit operands use mulldx.
  • Result is sign-extended to 64 bits. The 64-bit product fits into a 64-bit GPR without loss. Subsequent 32-bit consumers see RT[32:63] (the low 32 bits of the product); use mulhwx for the signed high 32 bits or mulhwux for the unsigned high 32 bits, computed in parallel without this instruction.
  • OE overflow test is 32-bit. XER[OV] is set iff the 64-bit signed product cannot be represented in 32 bits — iff RT[0:32] are not all equal (the product is not the sign extension of its low word). Canary's OE branch is XEINSTRNOTIMPLEMENTED().
  • CR0 compares only the low 32 bits in Canary. Canary stores the full 64-bit product of the sign-extended words, but f.UpdateCR(0, v) compares Truncate(v, INT32) with zero. The high 32 bits may be non-zero while the low 32 are zero, so Canary's CR0 can differ from spec's full 64-bit compare. It matters only for code that detects overflow through mullw.'s CR0 — rare.
  • Latency. On the Xenon, mullw has higher latency than add/sub; many hot inner loops avoid it by strength-reduction or shift-add chains. This is irrelevant for correctness but sometimes explains surprising instruction sequences in disassembly.
  • mulhwx — signed high 32 bits of the same 32×32 product.
  • mulhwux — unsigned high 32 bits of a 32×32 product.
  • mulli — D-form: RT ← (RA[32:63]) × SIMM (low 64 bits, signed).
  • mulldx, mulhdx, mulhdux — 64-bit multiplies (low/high, signed/unsigned).
  • divwx, divwux — 32-bit signed / unsigned division.

IBM Reference