Files
Sylpheed/tools/ppc-manual/vmx/vmaddfp.md
sim f3c512f2ab docs(ppc-manual): check every xenia-rs claim against Canary's source
The hand-written parts of the manual still described how the retired
xenia-rs interpreter behaved: its snapshots, Rust casts and helpers. Each of
those 490 statements is now either restated as what Canary's emitters and
x64 backend actually do (at the pinned canary_experimental commit), or
dropped where it only made sense for xenia-rs.

Checking them turned up claims that were wrong, not just outdated:

- VSCR[SAT] is never modelled in Canary (DID_SATURATE is a stub and mfvscr
  cannot see it); the pages said saturating ops set it stickily.
- Canary does not implement lswi/lswx/stswi/stswx, dcbi, mtfsb0/mtfsb1,
  vmsum*, vmhaddshs, vupkhpx/vupklpx, and most SPRs; pages described them
  as working.
- Traps evaluate TO in Canary; stvebx/stvehx/stvewx store one element, not
  16 bytes; mtmsrd writes only EE; fres/frsqrte/vrsqrtefp precision claims
  and the stfs "rounds under RN / sets FPSCR" claim contradicted the spec.
- Reservations are a 64 KiB block bitmap plus a value compare, not
  per-address tracking.

Claims that neither Canary's source nor a public spec settles are marked
unverified (NI at boot, vmaddcfp128 operand order, estimate bit-exactness).

Generated regions are untouched; re-running the generator changes nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 21:52:38 +02:00

182 lines
7.9 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# `vmaddfp` — Vector Multiply-Add Floating Point
> **Category:** [VMX (Altivec)](../categories/vmx.md) · **Form:** [VA](../forms/VA.md) · **Opcode:** `0x1000002e`
<!-- GENERATED: BEGIN -->
## Assembler Mnemonics
| Mnemonic | XML entry | Flags | Description |
| --- | --- | --- | --- |
| `vmaddfp` | `vmaddfp` | — | Vector Multiply-Add Floating Point |
| `vmaddfp128` | `vmaddfp128` | — | Vector128 Multiply Add Floating Point |
## Syntax
```asm
vmaddfp [VD], [VA], [VC], [VB]
vmaddfp128 [VD], [VA], [VB], [VD]
```
## Encoding
### `vmaddfp` — form `VA`
- **Opcode word:** `0x1000002e`
- **Primary opcode (bits 0–5):** `4`
- **Extended opcode:** `46`
- **Synchronising:** no
| Bits | Field | Meaning |
| --- | --- | --- |
| 0–5 | `OPCD` | primary opcode (4) |
| 6–10 | `VRT` | destination vector register |
| 11–15 | `VRA` | source A |
| 16–20 | `VRB` | source B |
| 21–25 | `VRC` | source C / shift |
| 26–31 | `XO` | extended opcode (6 bits) |
### `vmaddfp128` — form `VX128`
- **Opcode word:** `0x140000d0`
- **Primary opcode (bits 0–5):** `5`
- **Extended opcode:** `208`
- **Synchronising:** no
| Bits | Field | Meaning |
| --- | --- | --- |
| 0–5 | `OPCD` | primary opcode (4 or 5) |
| 6–10 | `VD128l` | destination low 5 bits |
| 11–15 | `VA128l` | source A low 5 bits |
| 16–20 | `VB128l` | source B low 5 bits |
| 21 | `VA128H` | source A high bit |
| 22 | `—` | reserved |
| 23–25 | `VC` | optional VC / XO sub-field |
| 26 | `VA128h` | source A middle bit |
| 27 | `—` | reserved |
| 28–29 | `VD128h` | destination high 2 bits |
| 30–31 | `VB128h` | source B high 2 bits |
## Operands
| Field | Role | Description |
| --- | --- | --- |
| `VA` | vmaddfp: read; vmaddfp128: read | Source A vector register. |
| `VC` | vmaddfp: read; vmaddfp128: read | Source C vector register / 3-bit selector. |
| `VB` | vmaddfp: read; vmaddfp128: read | Source B vector register. |
| `VD` | vmaddfp: write; vmaddfp128: write | Destination vector register. |
## Register Effects
### `vmaddfp`
- **Reads (always):** `VA`, `VC`, `VB`
- **Reads (conditional):** _none_
- **Writes (always):** `VD`
- **Writes (conditional):** _none_
### `vmaddfp128`
- **Reads (always):** `VA`, `VC`, `VB`
- **Reads (conditional):** _none_
- **Writes (always):** `VD`
- **Writes (conditional):** _none_
## Status-Register Effects
_No condition-register or status-register effects._
## Operation (pseudocode)
```
for each 32-bit float lane i in 0..3:
VD[i] <- (VA[i] * VC[i]) + VB[i]
```
## C Translation Example
```c
/* No hand-written C yet. Translate the Canary emitter snapshot */
/* under Implementation References; its HIR maps directly: */
/* f.LoadGPR(n) / f.StoreGPR(n, v) -> r[n] / r[n] = v */
/* f.LoadFPR / StoreFPR, f.LoadVR / StoreVR -> f[n], v[n] */
/* f.Load(ea, T), f.Store(ea, v) -> raw read / write; emitters */
/* wrap them in f.ByteSwap for the big-endian guest value */
/* f.UpdateCR(n, v) -> CR field n from v's LOW 32 BITS vs 0 */
/* f.LoadCA / f.StoreCA -> xer.CA; f.StoreSAT -> vscr.SAT */
/* i.XO.RA, i.D.DS, ... -> the bit-fields listed under Operands */
/* The Register Effects and Status-Register Effects tables above */
/* enumerate every side effect a faithful translation must emit. */
```
## Implementation References
**`vmaddfp`**
- Canary XML: [`tools/ppc-instructions.xml` — search for `mnem="vmaddfp"`](https://github.com/xenia-canary/xenia-canary/blob/f21ebd49e979e44f081f474df78c3fbfee9cb3f2/tools/ppc-instructions.xml)
- Canary emitter: [`src/xenia/cpu/ppc/ppc_emit_altivec.cc:795`](https://github.com/xenia-canary/xenia-canary/blob/f21ebd49e979e44f081f474df78c3fbfee9cb3f2/src/xenia/cpu/ppc/ppc_emit_altivec.cc#L795)
- Sylpheed opcode: [`crates/sylpheed-ppc/src/opcode.rs:348`](../../../crates/sylpheed-ppc/src/opcode.rs#L348)
- Sylpheed decoder: [`crates/sylpheed-ppc/src/decoder.rs:703`](../../../crates/sylpheed-ppc/src/decoder.rs#L703)
<details><summary>Canary emitter (frozen snapshot @ <code>f21ebd49e9</code>)</summary>
```cpp
int InstrEmit_vmaddfp(PPCHIRBuilder& f, const InstrData& i) {
// (VD) <- ((VA) * (VC)) + (VB)
return InstrEmit_vmaddfp_(f, i.VXA.VD, i.VXA.VA, i.VXA.VB, i.VXA.VC);
}
// ── delegates to (src/xenia/cpu/ppc/ppc_emit_altivec.cc:786) ──
int InstrEmit_vmaddfp_(PPCHIRBuilder& f, uint32_t vd, uint32_t va, uint32_t vb,
uint32_t vc) {
// POWER8 testing showed that vmaddfp flushes denormal inputs to zero
// regardless of NJM.
// (VD) <- ((VA) * (VC)) + (VB)
Value* v = f.MulAdd(f.LoadVR(va), f.LoadVR(vc), f.LoadVR(vb));
f.StoreVR(vd, v);
return 0;
}
```
</details>
**`vmaddfp128`**
- Canary XML: [`tools/ppc-instructions.xml` — search for `mnem="vmaddfp128"`](https://github.com/xenia-canary/xenia-canary/blob/f21ebd49e979e44f081f474df78c3fbfee9cb3f2/tools/ppc-instructions.xml)
- Canary emitter: [`src/xenia/cpu/ppc/ppc_emit_altivec.cc:799`](https://github.com/xenia-canary/xenia-canary/blob/f21ebd49e979e44f081f474df78c3fbfee9cb3f2/src/xenia/cpu/ppc/ppc_emit_altivec.cc#L799)
- Sylpheed opcode: [`crates/sylpheed-ppc/src/opcode.rs:349`](../../../crates/sylpheed-ppc/src/opcode.rs#L349)
- Sylpheed decoder: [`crates/sylpheed-ppc/src/decoder.rs:728`](../../../crates/sylpheed-ppc/src/decoder.rs#L728)
<details><summary>Canary emitter (frozen snapshot @ <code>f21ebd49e9</code>)</summary>
```cpp
int InstrEmit_vmaddfp128(PPCHIRBuilder& f, const InstrData& i) {
// (VD) <- ((VA) * (VB)) + (VD)
// NOTE: this resuses VD and swaps the arg order!
return InstrEmit_vmaddfp_(f, VX128_VD128, VX128_VA128, VX128_VD128,
VX128_VB128);
}
```
</details>
<!-- GENERATED: END -->
## Special Cases & Edge Conditions
- **Fused multiply-add: `VD = (VA * VC) + VB`** per word lane (single rounding). No intermediate rounding between the multiply and the add — this is critical for numerical accuracy in DSP filters and reduces error in dot products.
- **Big-endian word lanes.** Lane 0 is the most-significant word.
- **NaN propagation, ±∞ arithmetic.** Standard IEEE-754: any NaN input yields NaN; `(+∞ * 0)` yields NaN; the sum of `+∞` and `-∞` (e.g. `(+∞ * 1) + -∞`) yields NaN. No trap, no sticky bit.
- **`VSCR[NJ]` denormals.** With `NJ = 1` (Xenon default), denormal inputs and outputs are flushed to `±0`.
- **No `VSCR[SAT]` change, no XER change, no exceptions.**
- **VMX128 sibling has surprising operand layout — `VD` is also a source.** Canary's `vmaddfp128` passes `VD` as the addend, computing `VD = (VA * VB) + VD_prev` (its comment: "this resuses VD and swaps the arg order!"). The standard `vmaddfp` keeps the canonical 4-operand `VA, VC, VB → VD` shape. **This is a real difference in operand encoding** (VX128 form vs. VA-form) that compilers must respect — VMX128 sacrifices the third source register slot for the extra register-file bits.
- **Aliasing legal.** `vmaddfp v3, v3, v3, v3` works (squares + adds itself).
- **Common usage.** Per-lane polynomial evaluation, dot-product accumulation, any matrix multiply inner loop. Pair four `vmaddfp` instructions to do a 4×4 × 4-vec multiply.
## Related Instructions
- [`vnmsubfp`](vnmsubfp.md) — `−((VA * VC) − VB)`; fused negative-multiply-subtract.
- [`vaddfp`](vaddfp.md), [`vsubfp`](vsubfp.md) — plain float add / subtract.
- [`vmulfp128`](../vmx128/vmulfp128.md) — the VMX128-only `VA * VB` multiply; in standard Altivec, games use `vmaddfp v, va, vc, v0_zero` instead.
- [`vmaxfp`](vmaxfp.md), [`vminfp`](vminfp.md) — min / max for clamping.
- [`vrefp`](vrefp.md), [`vrsqrtefp`](vrsqrtefp.md) — reciprocal / inverse-sqrt estimates that often appear in the same FMA chain.
## IBM Reference
- [AIX 7.3 — `vmaddfp` (Vector Multiply-Add Floating Point)](https://www.ibm.com/docs/en/aix/7.3.0?topic=set-vmaddfp-vector-multiply-add-floating-point-instruction)
- [IBM AltiVec Technology Programmer's Interface Manual, Chapter 5 — Floating-Point Multiply-Add Family](https://www.nxp.com/docs/en/reference-manual/ALTIVECPIM.pdf)