The hand-written parts of the manual still described how the retired xenia-rs interpreter behaved: its snapshots, Rust casts and helpers. Each of those 490 statements is now either restated as what Canary's emitters and x64 backend actually do (at the pinned canary_experimental commit), or dropped where it only made sense for xenia-rs. Checking them turned up claims that were wrong, not just outdated: - VSCR[SAT] is never modelled in Canary (DID_SATURATE is a stub and mfvscr cannot see it); the pages said saturating ops set it stickily. - Canary does not implement lswi/lswx/stswi/stswx, dcbi, mtfsb0/mtfsb1, vmsum*, vmhaddshs, vupkhpx/vupklpx, and most SPRs; pages described them as working. - Traps evaluate TO in Canary; stvebx/stvehx/stvewx store one element, not 16 bytes; mtmsrd writes only EE; fres/frsqrte/vrsqrtefp precision claims and the stfs "rounds under RN / sets FPSCR" claim contradicted the spec. - Reservations are a 64 KiB block bitmap plus a value compare, not per-address tracking. Claims that neither Canary's source nor a public spec settles are marked unverified (NI at boot, vmaddcfp128 operand order, estimate bit-exactness). Generated regions are untouched; re-running the generator changes nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
128 lines
5.9 KiB
Markdown
128 lines
5.9 KiB
Markdown
# `fmaddx` — Floating Multiply-Add
|
||
|
||
> **Category:** [Floating-Point](../categories/fpu.md) · **Form:** [A](../forms/A.md) · **Opcode:** `0xfc00003a`
|
||
|
||
<!-- GENERATED: BEGIN -->
|
||
|
||
## Assembler Mnemonics
|
||
|
||
| Mnemonic | XML entry | Flags | Description |
|
||
| --- | --- | --- | --- |
|
||
| `fmadd` | `fmaddx` | — | Floating Multiply-Add |
|
||
| `fmadd.` | `fmaddx` | Rc=1 | Floating Multiply-Add |
|
||
|
||
## Syntax
|
||
|
||
```asm
|
||
fmadd[Rc] [FD], [FA], [FC], [FB]
|
||
```
|
||
|
||
## Encoding
|
||
|
||
### `fmaddx` — form `A`
|
||
|
||
- **Opcode word:** `0xfc00003a`
|
||
- **Primary opcode (bits 0–5):** `63`
|
||
- **Extended opcode:** `29`
|
||
- **Synchronising:** no
|
||
|
||
| Bits | Field | Meaning |
|
||
| --- | --- | --- |
|
||
| 0–5 | `OPCD` | primary opcode (59 or 63) |
|
||
| 6–10 | `FRT` | destination FPR |
|
||
| 11–15 | `FRA` | source A FPR |
|
||
| 16–20 | `FRB` | source B FPR |
|
||
| 21–25 | `FRC` | source C FPR (multiplier for madd-style ops) |
|
||
| 26–30 | `XO` | extended opcode (5 bits) |
|
||
| 31 | `Rc` | record-form flag (updates CR1) |
|
||
|
||
## Operands
|
||
|
||
| Field | Role | Description |
|
||
| --- | --- | --- |
|
||
| `FA` | fmaddx: read | Source A floating-point register (`fr0`–`fr31`). |
|
||
| `FC` | fmaddx: read | Source C floating-point register (for madd-style ops). |
|
||
| `FB` | fmaddx: read | Source B floating-point register. |
|
||
| `FD` | fmaddx: write | Destination floating-point register. |
|
||
| `CR` | fmaddx: write (conditional) | Condition-register update. When `Rc=1`, CR field 0 (or CR6 for vector compares, CR1 for FPU) is updated from the result. |
|
||
| `FPSCR` | fmaddx: write | Floating-Point Status and Control Register. |
|
||
|
||
## Register Effects
|
||
|
||
### `fmaddx`
|
||
|
||
- **Reads (always):** `FA`, `FC`, `FB`
|
||
- **Reads (conditional):** _none_
|
||
- **Writes (always):** `FD`, `FPSCR`
|
||
- **Writes (conditional):** `CR`
|
||
|
||
## Status-Register Effects
|
||
|
||
- `fmaddx`: **CR1** ← FPSCR[FX, FEX, VX, OX] when `Rc=1`.; **FPSCR** updated per IEEE-754 flags (FX, FEX, FPRF, FR, FI, exceptions).
|
||
|
||
## Operation (pseudocode)
|
||
|
||
```
|
||
FRT <- (FRA × FRC) + FRB
|
||
```
|
||
|
||
## C Translation Example
|
||
|
||
```c
|
||
/* No hand-written C yet. Translate the Canary emitter snapshot */
|
||
/* under Implementation References; its HIR maps directly: */
|
||
/* f.LoadGPR(n) / f.StoreGPR(n, v) -> r[n] / r[n] = v */
|
||
/* f.LoadFPR / StoreFPR, f.LoadVR / StoreVR -> f[n], v[n] */
|
||
/* f.Load(ea, T), f.Store(ea, v) -> raw read / write; emitters */
|
||
/* wrap them in f.ByteSwap for the big-endian guest value */
|
||
/* f.UpdateCR(n, v) -> CR field n from v's LOW 32 BITS vs 0 */
|
||
/* f.LoadCA / f.StoreCA -> xer.CA; f.StoreSAT -> vscr.SAT */
|
||
/* i.XO.RA, i.D.DS, ... -> the bit-fields listed under Operands */
|
||
/* The Register Effects and Status-Register Effects tables above */
|
||
/* enumerate every side effect a faithful translation must emit. */
|
||
```
|
||
|
||
## Implementation References
|
||
|
||
**`fmaddx`**
|
||
- Canary XML: [`tools/ppc-instructions.xml` — search for `mnem="fmaddx"`](https://github.com/xenia-canary/xenia-canary/blob/f21ebd49e979e44f081f474df78c3fbfee9cb3f2/tools/ppc-instructions.xml)
|
||
- Canary emitter: [`src/xenia/cpu/ppc/ppc_emit_fpu.cc:186`](https://github.com/xenia-canary/xenia-canary/blob/f21ebd49e979e44f081f474df78c3fbfee9cb3f2/src/xenia/cpu/ppc/ppc_emit_fpu.cc#L186)
|
||
- Sylpheed opcode: [`crates/sylpheed-ppc/src/opcode.rs:77`](../../../crates/sylpheed-ppc/src/opcode.rs#L77)
|
||
- Sylpheed decoder: [`crates/sylpheed-ppc/src/decoder.rs:1043`](../../../crates/sylpheed-ppc/src/decoder.rs#L1043)
|
||
<details><summary>Canary emitter (frozen snapshot @ <code>f21ebd49e9</code>)</summary>
|
||
|
||
```cpp
|
||
int InstrEmit_fmaddx(PPCHIRBuilder& f, const InstrData& i) {
|
||
return InstrEmit_fmadd(f, i, false);
|
||
}
|
||
```
|
||
</details>
|
||
|
||
<!-- GENERATED: END -->
|
||
|
||
## Special Cases & Edge Conditions
|
||
|
||
- **Single rounding step.** `fmadd` computes `(FRA × FRC) + FRB` with one IEEE-754 rounding at the end — strictly more accurate than separate multiply + add. Canary matches that only on hosts with FMA3, where it emits `vfmadd213sd`; without FMA3 it falls back to `vmulsd` + `vaddsd`, which rounds twice and can differ in the last bit.
|
||
- **Operand layout.** A-form: `FRT, FRA, FRC, FRB`. Note the assembler order — `FRC` (multiplier) comes before `FRB` (addend). Encoding bit fields are `FRA` (11–15), `FRB` (16–20), `FRC` (21–25).
|
||
- **Invalid operations.** `0×∞ + finite` → `VXIMZ`; `∞×x + ∓∞` (after multiplication produces ±∞ that opposes addend sign) → `VXISI`. Quiet NaN result with `FPSCR[VX, FX]` set.
|
||
- **FPSCR side effects.** Hardware updates `FPRF`, `FR`, `FI`, `FX`, `OX`, `UX`, `XX`, `VXIMZ`, `VXISI`, `VXSNAN`. Canary does not update FPSCR (`UpdateFPSCR` is a stub).
|
||
- **`Rc=1` (`fmadd.`)** copies `FPSCR[FX, FEX, VX, OX]` into CR1.
|
||
- **NaN propagation.** Quiet-NaN result for any NaN operand; signalling NaNs are quietened.
|
||
- **Use case.** Dot products, polynomial evaluation (Horner's method), matrix multiplies, Newton-Raphson divide/sqrt refinement. Hot-path PPC code is dense with `fmadd`.
|
||
- **Denormal flush.** That Xenon boots with `FPSCR[NI]=1` is unverified. Canary starts every guest thread in IEEE mode (MXCSR `0x1F80`; its init comment flags the startup state as unchecked) and turns on the host's flush-to-zero (`MXCSR.FZ`) only when the guest sets `NI` through `mtfsf`/`mtfsfi`.
|
||
|
||
## Related Instructions
|
||
|
||
- [`fmaddsx`](fmaddsx.md) — single-precision sibling.
|
||
- [`fmsubx`](fmsubx.md), [`fnmaddx`](fnmaddx.md), [`fnmsubx`](fnmsubx.md) — the other three fused multiply-add variants:
|
||
- `fmsub` = `(A×C) − B`
|
||
- `fnmadd` = `−((A×C) + B)`
|
||
- `fnmsub` = `−((A×C) − B)`
|
||
- [`fmulx`](fmulx.md), [`faddx`](faddx.md) — non-fused decomposition (two rounding steps; less precise).
|
||
- [`fresx`](fresx.md), [`frsqrtex`](frsqrtex.md) — reciprocal helpers refined by `fmadd`/`fnmsub`.
|
||
|
||
## IBM Reference
|
||
|
||
- [AIX 7.3 — `fmadd` (Floating Multiply-Add)](https://www.ibm.com/docs/en/aix/7.3.0?topic=set-fma-fmadd-floating-multiply-add-instruction)
|
||
- [PowerISA v2.07B, Book I, Chapter 4 — Floating-Point Processor](https://openpowerfoundation.org/specifications/isa/) (single-rounding fused multiply-add definition).
|