Files
Sylpheed/tools/ppc-manual/memory/lmw.md
sim f3c512f2ab docs(ppc-manual): check every xenia-rs claim against Canary's source
The hand-written parts of the manual still described how the retired
xenia-rs interpreter behaved: its snapshots, Rust casts and helpers. Each of
those 490 statements is now either restated as what Canary's emitters and
x64 backend actually do (at the pinned canary_experimental commit), or
dropped where it only made sense for xenia-rs.

Checking them turned up claims that were wrong, not just outdated:

- VSCR[SAT] is never modelled in Canary (DID_SATURATE is a stub and mfvscr
  cannot see it); the pages said saturating ops set it stickily.
- Canary does not implement lswi/lswx/stswi/stswx, dcbi, mtfsb0/mtfsb1,
  vmsum*, vmhaddshs, vupkhpx/vupklpx, and most SPRs; pages described them
  as working.
- Traps evaluate TO in Canary; stvebx/stvehx/stvewx store one element, not
  16 bytes; mtmsrd writes only EE; fres/frsqrte/vrsqrtefp precision claims
  and the stfs "rounds under RN / sets FPSCR" claim contradicted the spec.
- Reservations are a 64 KiB block bitmap plus a value compare, not
  per-address tracking.

Claims that neither Canary's source nor a public spec settles are marked
unverified (NI at boot, vmaddcfp128 operand order, estimate bit-exactness).

Generated regions are untouched; re-running the generator changes nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 21:52:38 +02:00

5.6 KiB
Raw Permalink Blame History

lmw — Load Multiple Word

Category: Memory · Form: D · Opcode: 0xb8000000

Assembler Mnemonics

Mnemonic XML entry Flags Description
lmw lmw Load Multiple Word

Syntax

(no disassembly template)

Encoding

lmw — form D

  • Opcode word: 0xb8000000
  • Primary opcode (bits 05): 46
  • Extended opcode:
  • Synchronising: no
Bits Field Meaning
05 OPCD primary opcode
610 RT destination GPR (or RS when storing)
1115 RA source GPR (0 ⇒ literal 0 for RA0 forms)
1631 D/SI/UI 16-bit signed or unsigned immediate

Operands

Field Role Description

Register Effects

lmw

  • Reads (always): none
  • Reads (conditional): none
  • Writes (always): none
  • Writes (conditional): none

Status-Register Effects

No condition-register or status-register effects.

Operation (pseudocode)

; No hand-written pseudocode for this instruction yet.
; The authoritative semantics are the Canary emitter snapshot under
; Implementation References; about half of Canary's emitters open
; with the PPC-style definition as a comment (`RD <- (RA) + (RB)`).
; Every side effect is also enumerated in the Register Effects and
; Status-Register Effects tables above.

C Translation Example

/* No hand-written C yet. Translate the Canary emitter snapshot   */
/* under Implementation References; its HIR maps directly:        */
/*   f.LoadGPR(n) / f.StoreGPR(n, v)  -> r[n] / r[n] = v          */
/*   f.LoadFPR / StoreFPR, f.LoadVR / StoreVR -> f[n], v[n]        */
/*   f.Load(ea, T), f.Store(ea, v) -> raw read / write; emitters   */
/*     wrap them in f.ByteSwap for the big-endian guest value      */
/*   f.UpdateCR(n, v)  -> CR field n from v's LOW 32 BITS vs 0     */
/*   f.LoadCA / f.StoreCA -> xer.CA;  f.StoreSAT -> vscr.SAT       */
/*   i.XO.RA, i.D.DS, ... -> the bit-fields listed under Operands  */
/* The Register Effects and Status-Register Effects tables above  */
/* enumerate every side effect a faithful translation must emit.  */

Implementation References

lmw

Canary emitter (frozen snapshot @ f21ebd49e9)
int InstrEmit_lmw(PPCHIRBuilder& f, const InstrData& i) {
  Value* b;
  if (i.D.RA == 0) {
    b = f.LoadZeroInt64();
  } else {
    b = f.LoadGPR(i.D.RA);
  }

  for (uint32_t j = 0; j < 32 - i.D.RT; ++j) {
    if (i.D.RT + j == i.D.RA) {
      continue;
    }
    Value* offset = f.LoadConstantInt64(XEEXTS16(i.D.DS) + j * 4);
    Value* rt = f.ZeroExtend(f.ByteSwap(f.LoadOffset(b, offset, INT32_TYPE)),
                             INT64_TYPE);
    f.StoreGPR(i.D.RT + j, rt);
  }
  return 0;
}

Special Cases & Edge Conditions

  • Bulk register restore. Loads (32 - RT) consecutive 32-bit words starting at EA into RT, RT+1, …, r31. Used by AIX/PowerPC ABI prologues/epilogues to restore non-volatile GPRs in one instruction. Modern compilers prefer multiple lwz for scheduling; lmw survives in older code and hand-rolled context-switch routines.
  • Loop bound from encoding. Canary loops over r(RT)..r31 (j < 32 - RT), matching IBM's "load until r31 inclusive" semantic. With RT = 28, four registers (r28..r31) are loaded. ⚠️ If RA falls in that range (an invalid form), Canary skips the load into RA, so the base register keeps its value.
  • Each word is zero-extended. Like lwz, every loaded 32-bit word zero-extends into the destination's 64-bit GPR. The high 32 bits of each r[k] become zero.
  • Big-endian read. Word at EA goes to r[RT], word at EA+4 goes to r[RT+1], etc. Each word is itself loaded most-significant-byte-first.
  • RA0 semantics. When RA = 0, base is literal zero. Useful for absolute-address restoration.
  • Invalid forms. AIX docs declare it invalid for RA to be in the destination range [RT, 31] — a load could overwrite the base register mid-sequence. Canary skips the load into RA in that case, so the base register keeps its value.
  • Alignment. PowerISA requires word-aligned EA; an unaligned lmw may raise an alignment exception on real hardware. Canary does not check.
  • Performance trap. On modern PowerPC implementations lmw is microcoded — slower than the equivalent sequence of lwz. Compilers avoid it.
  • stmw — symmetric "store multiple words" (the matching epilogue/prologue partner).
  • lwz, lwzx — single-word loads; the modern preferred form.
  • lswi, lswx — load string (byte-granular bulk transfer).

IBM Reference