Files
Sylpheed/tools/ppc-manual/alu/mullwx.md
sim dedcf37867 docs(ppc-manual): quote Canary and our own decoder, not the retired xenia-rs
The generator had not been able to run correctly since the manual moved into
`tools/ppc-manual/`: it computed the repository root as `HERE.parent.parent`,
which now names `tools/`, so the XML, Canary's emitters and xenia-rs all stopped
resolving — silently, because both scrapers skipped what they could not find.
Every page's references had been pointing at paths that exist nowhere.

What each source contributed, measured on the 350 pages before this change:

  Operation (pseudocode)  251 pages: fixed boilerplate "derives from the xenia-rs
                          interpreter"; 99 carry real hand-written seeds
  C translation           337 pages: the same kind of boilerplate
  xenia-rs snapshot       336 pages: the interpreter arm, pasted in — the only
                          per-instruction semantics on unseeded pages
  links                   xenia-rs opcode/decoder/interpreter + Canary emitter

Now:

  * semantics come from **Xenia Canary**, the reference emulator, read through
    `git show` at a pinned upstream commit (`origin/canary_experimental`,
    f21ebd49e9). Not our checkout: it carries instrumentation and lacked
    upstream's `mcrf` fix, so it would have published probes and a wrong `mcrf`.
    Each page embeds the emitter (`InstrEmit_<mnem>`), and for the 128 pure
    one-line delegations also the helper that holds the semantics.
  * decode references point at `crates/sylpheed-ppc` — the decoder that
    produces `sylpheed.db` — as in-repo relative links.
  * the boilerplate now says what is true, and the C translation guide maps
    Canary's actual HIR calls, checked against `ppc_hir_builder.h` (including
    that `UpdateCR(n, v)` truncates to 32 bits).
  * `rust_scraper.py` -> `decoder_scraper.py` (interpreter half dropped);
    missing sources are now errors, not empty results.

Verified:

  consistency checks        455 XML entries, 350 families, 598 index keys
  hand-written tails        386/386 byte-identical after regeneration
  xenia-rs in generated     0
  pages with a snapshot     349/350 (was 336) — `dcbi` has no Canary emitter at all
  in-repo decoder links     910/910 resolve to a line holding the identifier
  emitter boundaries        brace counter == column-0 `}` rule on 521/521;
                            preprocessor model unit-tested (#if 0/#else/#elif)
  idempotency               re-run: 0 pages updated, 0 working-tree changes

Hand-written notes (outside the generated regions) are not rewritten here:

  * 110 links into `../../xenia-rs/...` were dead; they now point at the file in
    the archived repository (git.mc02.dev/fabi/xenia-rs @ 8401d4d). Line anchors
    were dropped because the notes predate that commit — 0 of 441 old line
    ranges match it — and a precise-looking wrong anchor is worse than none. The
    link text, which carries the author's line numbers, is unchanged.
  * 140 prose claims about xenia-rs's behaviour remain. 23 are verified to hold
    for Canary too (the 32-bit CR0 truncation, OE left unimplemented); the other
    114 need checking one by one, and some invert — e.g. `divdx` notes a correct
    64-bit CR0 update in xenia-rs where Canary's `UpdateCR` truncates. Left for
    a deliberate pass rather than a blind substitution.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 20:34:34 +02:00

6.6 KiB
Raw Blame History

mullwx — Multiply Low Word

Category: Integer ALU · Form: XO · Opcode: 0x7c0001d6

Assembler Mnemonics

Mnemonic XML entry Flags Description
mullw mullwx Multiply Low Word
mullwo mullwx OE=1 Multiply Low Word
mullw. mullwx Rc=1 Multiply Low Word
mullwo. mullwx OE=1, Rc=1 Multiply Low Word

Syntax

mullw[OE][Rc] [RD], [RA], [RB]

Encoding

mullwx — form XO

  • Opcode word: 0x7c0001d6
  • Primary opcode (bits 05): 31
  • Extended opcode: 235
  • Synchronising: no
Bits Field Meaning
05 OPCD primary opcode (31)
610 RT destination GPR
1115 RA source A
1620 RB source B
21 OE overflow-enable flag
2230 XO extended opcode (9 bits)
31 Rc record-form flag

Operands

Field Role Description
RA mullwx: read Source GPR (r0r31).
RB mullwx: read Source GPR.
RD mullwx: write Destination GPR.
CR mullwx: write (conditional) Condition-register update. When Rc=1, CR field 0 (or CR6 for vector compares, CR1 for FPU) is updated from the result.
OE mullwx: write (conditional) Overflow-enable bit. When 1, the instruction updates XER[OV] and stickies XER[SO] on signed overflow.

Register Effects

mullwx

  • Reads (always): RA, RB
  • Reads (conditional): none
  • Writes (always): RD
  • Writes (conditional): CR, OE

Status-Register Effects

  • mullwx: CR0 ← signed-compare(result, 0) with SO ← XER[SO], when Rc=1.; XER[OV] ← signed-overflow(result); XER[SO] stickies, when OE=1.

Operation (pseudocode)

RT <- ((RA)[32:63]) * ((RB)[32:63])    ; signed 32×32 → 64

C Translation Example

/* No hand-written C yet. Translate the Canary emitter snapshot   */
/* under Implementation References; its HIR maps directly:        */
/*   f.LoadGPR(n) / f.StoreGPR(n, v)  -> r[n] / r[n] = v          */
/*   f.LoadFPR / StoreFPR, f.LoadVR / StoreVR -> f[n], v[n]        */
/*   f.Load(ea, T), f.Store(ea, v) -> raw read / write; emitters   */
/*     wrap them in f.ByteSwap for the big-endian guest value      */
/*   f.UpdateCR(n, v)  -> CR field n from v's LOW 32 BITS vs 0     */
/*   f.LoadCA / f.StoreCA -> xer.CA;  f.StoreSAT -> vscr.SAT       */
/*   i.XO.RA, i.D.DS, ... -> the bit-fields listed under Operands  */
/* The Register Effects and Status-Register Effects tables above  */
/* enumerate every side effect a faithful translation must emit.  */

Implementation References

mullwx

Canary emitter (frozen snapshot @ f21ebd49e9)
int InstrEmit_mullwx(PPCHIRBuilder& f, const InstrData& i) {
  // RT <- (RA)[32:63] × (RB)[32:63]
  if (i.XO.OE) {
    // With XER update.
    XEINSTRNOTIMPLEMENTED();
  }
  Value* v = f.Mul(
      f.SignExtend(f.Truncate(f.LoadGPR(i.XO.RA), INT32_TYPE), INT64_TYPE),
      f.SignExtend(f.Truncate(f.LoadGPR(i.XO.RB), INT32_TYPE), INT64_TYPE));
  f.StoreGPR(i.XO.RT, v);
  if (i.XO.Rc) {
    f.UpdateCR(0, v);
  }
  return 0;
}

Extended Pseudocode

prod64 <- sign_extend_32_to_64((RA)[32:63]) *s sign_extend_32_to_64((RB)[32:63])
RT <- prod64                                         ; 64-bit result
if OE then
    XER[OV] <- (prod64 ≠ sign_extend_32_to_64(prod64[32:63]))  ; set when product doesn't fit in 32 bits
    XER[SO] <- XER[SO] | XER[OV]
if Rc then
    CR0 <- signed_compare(RT, 0) || XER[SO]

Special Cases & Edge Conditions

  • Inputs are the low 32 bits. mullw only looks at RA[32:63] and RB[32:63]; the high 32 bits of each source are ignored. This is a 32-bit × 32-bit → 64-bit signed multiply. For full 64-bit operands use mulldx.
  • Result is sign-extended to 64 bits. The 64-bit product fits into a 64-bit GPR without loss. Subsequent 32-bit consumers see RT[32:63] (the low 32 bits of the product); use mulhwx for the signed high 32 bits or mulhwux for the unsigned high 32 bits, computed in parallel without this instruction.
  • OE overflow test is 32-bit. XER[OV] is set iff the 64-bit signed product cannot be represented in 32 bits — equivalently, iff RT[32] ≠ RT[33] = … = RT[63] (sign bit disagrees with the next 32 bits). Xenia-rs does not implement this; OE on mullwo is a no-op in the interpreter.
  • Xenia-rs CR0 update bug footprint. The interpreter computes CR0 from result as i32 as i64 — the low 32 bits sign-extended. For a 32×32→64 multiply the high 32 bits may be non-zero even when the low 32 bits are zero, so xenia's CR0 can differ from the spec's (which compares the full 64-bit product to zero). In practice this matters only for code that relies on mullw. to detect overflow via CR0 — extremely rare.
  • Latency. On the Xenon, mullw has higher latency than add/sub; many hot inner loops avoid it by strength-reduction or shift-add chains. This is irrelevant for correctness but sometimes explains surprising instruction sequences in disassembly.
  • mulhwx — signed high 32 bits of the same 32×32 product.
  • mulhwux — unsigned high 32 bits of a 32×32 product.
  • mulli — D-form: RT ← (RA[32:63]) × SIMM (low 64 bits, signed).
  • mulldx, mulhdx, mulhdux — 64-bit multiplies (low/high, signed/unsigned).
  • divwx, divwux — 32-bit signed / unsigned division.

IBM Reference