The hand-written parts of the manual still described how the retired xenia-rs interpreter behaved: its snapshots, Rust casts and helpers. Each of those 490 statements is now either restated as what Canary's emitters and x64 backend actually do (at the pinned canary_experimental commit), or dropped where it only made sense for xenia-rs. Checking them turned up claims that were wrong, not just outdated: - VSCR[SAT] is never modelled in Canary (DID_SATURATE is a stub and mfvscr cannot see it); the pages said saturating ops set it stickily. - Canary does not implement lswi/lswx/stswi/stswx, dcbi, mtfsb0/mtfsb1, vmsum*, vmhaddshs, vupkhpx/vupklpx, and most SPRs; pages described them as working. - Traps evaluate TO in Canary; stvebx/stvehx/stvewx store one element, not 16 bytes; mtmsrd writes only EE; fres/frsqrte/vrsqrtefp precision claims and the stfs "rounds under RN / sets FPSCR" claim contradicted the spec. - Reservations are a 64 KiB block bitmap plus a value compare, not per-address tracking. Claims that neither Canary's source nor a public spec settles are marked unverified (NI at boot, vmaddcfp128 operand order, estimate bit-exactness). Generated regions are untouched; re-running the generator changes nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
176 lines
7.4 KiB
Markdown
176 lines
7.4 KiB
Markdown
# `stvx` — Store Vector Indexed
|
||
|
||
> **Category:** [Memory](../categories/memory.md) · **Form:** [X](../forms/X.md) · **Opcode:** `0x7c0001ce`
|
||
|
||
<!-- GENERATED: BEGIN -->
|
||
|
||
## Assembler Mnemonics
|
||
|
||
| Mnemonic | XML entry | Flags | Description |
|
||
| --- | --- | --- | --- |
|
||
| `stvx` | `stvx` | — | Store Vector Indexed |
|
||
| `stvx128` | `stvx128` | — | Store Vector Indexed 128 |
|
||
|
||
## Syntax
|
||
|
||
```asm
|
||
stvx [VS], [RA0], [RB]
|
||
stvx128 [VS], [RA0], [RB]
|
||
```
|
||
|
||
## Encoding
|
||
|
||
### `stvx` — form `X`
|
||
|
||
- **Opcode word:** `0x7c0001ce`
|
||
- **Primary opcode (bits 0–5):** `31`
|
||
- **Extended opcode:** `231`
|
||
- **Synchronising:** no
|
||
|
||
| Bits | Field | Meaning |
|
||
| --- | --- | --- |
|
||
| 0–5 | `OPCD` | primary opcode |
|
||
| 6–10 | `RT/FRT/VRT` | destination |
|
||
| 11–15 | `RA/FRA/VRA` | source A |
|
||
| 16–20 | `RB/FRB/VRB` | source B |
|
||
| 21–30 | `XO` | extended opcode (10 bits) |
|
||
| 31 | `Rc` | record-form flag |
|
||
|
||
### `stvx128` — form `VX128_1`
|
||
|
||
- **Opcode word:** `0x100001c3`
|
||
- **Primary opcode (bits 0–5):** `4`
|
||
- **Extended opcode:** `451`
|
||
- **Synchronising:** no
|
||
|
||
| Bits | Field | Meaning |
|
||
| --- | --- | --- |
|
||
| 0–5 | `OPCD` | primary opcode (4) |
|
||
| 6–10 | `VD128l` | destination low 5 bits |
|
||
| 11–15 | `RA` | address register |
|
||
| 16–20 | `RB` | offset register |
|
||
| 21–27 | `XO` | extended opcode |
|
||
| 28–29 | `VD128h` | destination high 2 bits |
|
||
| 30–31 | `—` | reserved |
|
||
|
||
## Operands
|
||
|
||
| Field | Role | Description |
|
||
| --- | --- | --- |
|
||
| `VS` | stvx: read; stvx128: read | Source vector register (alias for VD on stores). |
|
||
| `RA0` | stvx: read; stvx128: read | Source GPR; when the encoded register number is 0 the operand is the literal 64-bit zero, **not** `r0`. |
|
||
| `RB` | stvx: read; stvx128: read | Source GPR. |
|
||
|
||
## Register Effects
|
||
|
||
### `stvx`
|
||
|
||
- **Reads (always):** `VS`, `RA0`, `RB`
|
||
- **Reads (conditional):** _none_
|
||
- **Writes (always):** _none_
|
||
- **Writes (conditional):** _none_
|
||
|
||
### `stvx128`
|
||
|
||
- **Reads (always):** `VS`, `RA0`, `RB`
|
||
- **Reads (conditional):** _none_
|
||
- **Writes (always):** _none_
|
||
- **Writes (conditional):** _none_
|
||
|
||
## Status-Register Effects
|
||
|
||
_No condition-register or status-register effects._
|
||
|
||
## Operation (pseudocode)
|
||
|
||
```
|
||
EA <- ((RA|0) + (RB)) & ~0xF ; align to 16
|
||
MEM(EA, 16) <- byteswap(VS)
|
||
```
|
||
|
||
## C Translation Example
|
||
|
||
```c
|
||
/* stvx VS, RA, RB — 16-byte aligned store of a vector register */
|
||
uint64_t base = (insn.RA == 0) ? 0 : r[insn.RA];
|
||
uint32_t ea = (uint32_t)((base + r[insn.RB]) & ~(uint64_t)0xF);
|
||
mem_write_vec128_be(ea, v[insn.VS]);
|
||
```
|
||
|
||
## Implementation References
|
||
|
||
**`stvx`**
|
||
- Canary XML: [`tools/ppc-instructions.xml` — search for `mnem="stvx"`](https://github.com/xenia-canary/xenia-canary/blob/f21ebd49e979e44f081f474df78c3fbfee9cb3f2/tools/ppc-instructions.xml)
|
||
- Canary emitter: [`src/xenia/cpu/ppc/ppc_emit_altivec.cc:193`](https://github.com/xenia-canary/xenia-canary/blob/f21ebd49e979e44f081f474df78c3fbfee9cb3f2/src/xenia/cpu/ppc/ppc_emit_altivec.cc#L193)
|
||
- Sylpheed opcode: [`crates/sylpheed-ppc/src/opcode.rs:269`](../../../crates/sylpheed-ppc/src/opcode.rs#L269)
|
||
- Sylpheed decoder: [`crates/sylpheed-ppc/src/decoder.rs:906`](../../../crates/sylpheed-ppc/src/decoder.rs#L906)
|
||
<details><summary>Canary emitter (frozen snapshot @ <code>f21ebd49e9</code>)</summary>
|
||
|
||
```cpp
|
||
int InstrEmit_stvx(PPCHIRBuilder& f, const InstrData& i) {
|
||
return InstrEmit_stvx_(f, i, i.X.RT, i.X.RA, i.X.RB);
|
||
}
|
||
|
||
// ── delegates to (src/xenia/cpu/ppc/ppc_emit_altivec.cc:187) ──
|
||
int InstrEmit_stvx_(PPCHIRBuilder& f, const InstrData& i, uint32_t vd,
|
||
uint32_t ra, uint32_t rb) {
|
||
Value* ea = f.And(CalculateEA_0(f, ra, rb), f.LoadConstantUint64(~0xFull));
|
||
f.Store(ea, f.ByteSwap(f.LoadVR(vd)));
|
||
return 0;
|
||
}
|
||
```
|
||
</details>
|
||
|
||
**`stvx128`**
|
||
- Canary XML: [`tools/ppc-instructions.xml` — search for `mnem="stvx128"`](https://github.com/xenia-canary/xenia-canary/blob/f21ebd49e979e44f081f474df78c3fbfee9cb3f2/tools/ppc-instructions.xml)
|
||
- Canary emitter: [`src/xenia/cpu/ppc/ppc_emit_altivec.cc:196`](https://github.com/xenia-canary/xenia-canary/blob/f21ebd49e979e44f081f474df78c3fbfee9cb3f2/src/xenia/cpu/ppc/ppc_emit_altivec.cc#L196)
|
||
- Sylpheed opcode: [`crates/sylpheed-ppc/src/opcode.rs:270`](../../../crates/sylpheed-ppc/src/opcode.rs#L270)
|
||
- Sylpheed decoder: [`crates/sylpheed-ppc/src/decoder.rs:532`](../../../crates/sylpheed-ppc/src/decoder.rs#L532)
|
||
<details><summary>Canary emitter (frozen snapshot @ <code>f21ebd49e9</code>)</summary>
|
||
|
||
```cpp
|
||
int InstrEmit_stvx128(PPCHIRBuilder& f, const InstrData& i) {
|
||
return InstrEmit_stvx_(f, i, VX128_1_VD128, i.VX128_1.RA, i.VX128_1.RB);
|
||
}
|
||
|
||
// ── delegates to (src/xenia/cpu/ppc/ppc_emit_altivec.cc:187) ──
|
||
int InstrEmit_stvx_(PPCHIRBuilder& f, const InstrData& i, uint32_t vd,
|
||
uint32_t ra, uint32_t rb) {
|
||
Value* ea = f.And(CalculateEA_0(f, ra, rb), f.LoadConstantUint64(~0xFull));
|
||
f.Store(ea, f.ByteSwap(f.LoadVR(vd)));
|
||
return 0;
|
||
}
|
||
```
|
||
</details>
|
||
|
||
<!-- GENERATED: END -->
|
||
|
||
## Extended Pseudocode
|
||
|
||
```
|
||
EA <- ((RA|0) + (RB)) & ~0xF ; force 16-byte alignment
|
||
MEM(EA, 16) <- byte_order_adjusted(VS) ; lane 0 at EA, lane 15 at EA+15
|
||
```
|
||
|
||
## Special Cases & Edge Conditions
|
||
|
||
- **Alignment is forced, not checked.** The low four bits of the effective address are **cleared** before the store — alignment violations silently corrupt adjacent data rather than trap. This differs from scalar `stw` (no alignment enforcement) and from `stvewx` (which stores only one element and keeps the exact EA).
|
||
- **Big-endian lane layout.** Vector lane 0 (the most-significant bytes of the 128-bit register) lives at the lowest address; lane 15 at `EA + 15`. On little-endian hosts the whole 16-byte block is byte-swapped at the memory boundary so the PowerPC-visible layout is preserved. Canary rounds `EA` down to a 16-byte boundary and stores `ByteSwap(VR)`.
|
||
- **`RA0` semantics.** When `RA = 0` the base is the literal zero — just like scalar loads/stores. Combined with the alignment mask this lets `stvx VS, 0, RB` store to address `RB & ~0xF`.
|
||
- **No update form.** Unlike scalar stores, VMX stores have no `u` variant that post-writes the base. Use [`stvxl`](stvxl.md) for the cache-hint variant (suggests "last" — the line is not expected to be reused soon).
|
||
- **VMX128 sibling (`stvx128`).** Identical semantics; the only difference is the operand encoding. VMX128 uses a 7-bit register index split across three non-contiguous bit fields (`VS128l ‖ VS128h`) so it can address `v0..v127` instead of the 32-register Altivec space. All alignment, byte-order and `RA0` rules are the same.
|
||
- **Read-before-write.** The 16-byte write occurs as one conceptual store; subsequent loads from the same address observe the complete new value. There's no split-transaction window visible to software.
|
||
|
||
## Related Instructions
|
||
|
||
- [`lvx`](lvx.md), [`lvx128`](lvx.md) — the load counterparts.
|
||
- [`stvxl`](stvxl.md), [`stvxl128`](stvxl.md) — cache-hint "last-use" variants.
|
||
- [`stvebx`](stvebx.md) / [`stvehx`](stvehx.md) / [`stvewx`](stvewx.md) — store single element (byte / half / word) at the exact (unaligned) address.
|
||
- [`stvlx`](stvlx.md) / [`stvrx`](stvrx.md) — store-left / store-right for unaligned vector I/O.
|
||
- [`dcbz`](dcbz.md) — zero a cache line; often paired with `stvx` in block-fill idioms.
|
||
|
||
## IBM Reference
|
||
|
||
- [AIX 7.3 — `stvx` (Store Vector Indexed)](https://www.ibm.com/docs/en/aix/7.3.0?topic=set-stvx-store-vector-indexed-instruction)
|
||
- PowerISA Book I, Vector facility (VMX / AltiVec). Xbox 360 VMX128 is Microsoft-documented in the XDK; Canary's `tools/ppc-instructions.xml` captures the deltas.
|