The generator had not been able to run correctly since the manual moved into
`tools/ppc-manual/`: it computed the repository root as `HERE.parent.parent`,
which now names `tools/`, so the XML, Canary's emitters and xenia-rs all stopped
resolving — silently, because both scrapers skipped what they could not find.
Every page's references had been pointing at paths that exist nowhere.
What each source contributed, measured on the 350 pages before this change:
Operation (pseudocode) 251 pages: fixed boilerplate "derives from the xenia-rs
interpreter"; 99 carry real hand-written seeds
C translation 337 pages: the same kind of boilerplate
xenia-rs snapshot 336 pages: the interpreter arm, pasted in — the only
per-instruction semantics on unseeded pages
links xenia-rs opcode/decoder/interpreter + Canary emitter
Now:
* semantics come from **Xenia Canary**, the reference emulator, read through
`git show` at a pinned upstream commit (`origin/canary_experimental`,
f21ebd49e9). Not our checkout: it carries instrumentation and lacked
upstream's `mcrf` fix, so it would have published probes and a wrong `mcrf`.
Each page embeds the emitter (`InstrEmit_<mnem>`), and for the 128 pure
one-line delegations also the helper that holds the semantics.
* decode references point at `crates/sylpheed-ppc` — the decoder that
produces `sylpheed.db` — as in-repo relative links.
* the boilerplate now says what is true, and the C translation guide maps
Canary's actual HIR calls, checked against `ppc_hir_builder.h` (including
that `UpdateCR(n, v)` truncates to 32 bits).
* `rust_scraper.py` -> `decoder_scraper.py` (interpreter half dropped);
missing sources are now errors, not empty results.
Verified:
consistency checks 455 XML entries, 350 families, 598 index keys
hand-written tails 386/386 byte-identical after regeneration
xenia-rs in generated 0
pages with a snapshot 349/350 (was 336) — `dcbi` has no Canary emitter at all
in-repo decoder links 910/910 resolve to a line holding the identifier
emitter boundaries brace counter == column-0 `}` rule on 521/521;
preprocessor model unit-tested (#if 0/#else/#elif)
idempotency re-run: 0 pages updated, 0 working-tree changes
Hand-written notes (outside the generated regions) are not rewritten here:
* 110 links into `../../xenia-rs/...` were dead; they now point at the file in
the archived repository (git.mc02.dev/fabi/xenia-rs @ 8401d4d). Line anchors
were dropped because the notes predate that commit — 0 of 441 old line
ranges match it — and a precise-looking wrong anchor is worse than none. The
link text, which carries the author's line numbers, is unchanged.
* 140 prose claims about xenia-rs's behaviour remain. 23 are verified to hold
for Canary too (the 32-bit CR0 truncation, OE left unimplemented); the other
114 need checking one by one, and some invert — e.g. `divdx` notes a correct
64-bit CR0 update in xenia-rs where Canary's `UpdateCR` truncates. Left for
a deliberate pass rather than a blind substitution.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
7.8 KiB
7.8 KiB
lvsl — Load Vector for Shift Left Indexed
Category: VMX (Altivec) · Form: X · Opcode:
0x7c00000c
Assembler Mnemonics
| Mnemonic | XML entry | Flags | Description |
|---|---|---|---|
lvsl |
lvsl |
— | Load Vector for Shift Left Indexed |
lvsl128 |
lvsl128 |
— | Load Vector for Shift Left Indexed 128 |
Syntax
lvsl [VD], [RA0], [RB]
lvsl128 [VD], [RA0], [RB]
Encoding
lvsl — form X
- Opcode word:
0x7c00000c - Primary opcode (bits 0–5):
31 - Extended opcode:
6 - Synchronising: no
| Bits | Field | Meaning |
|---|---|---|
| 0–5 | OPCD |
primary opcode |
| 6–10 | RT/FRT/VRT |
destination |
| 11–15 | RA/FRA/VRA |
source A |
| 16–20 | RB/FRB/VRB |
source B |
| 21–30 | XO |
extended opcode (10 bits) |
| 31 | Rc |
record-form flag |
lvsl128 — form VX128_1
- Opcode word:
0x10000003 - Primary opcode (bits 0–5):
4 - Extended opcode:
3 - Synchronising: no
| Bits | Field | Meaning |
|---|---|---|
| 0–5 | OPCD |
primary opcode (4) |
| 6–10 | VD128l |
destination low 5 bits |
| 11–15 | RA |
address register |
| 16–20 | RB |
offset register |
| 21–27 | XO |
extended opcode |
| 28–29 | VD128h |
destination high 2 bits |
| 30–31 | — |
reserved |
Operands
| Field | Role | Description |
|---|---|---|
RA0 |
lvsl: read; lvsl128: read | Source GPR; when the encoded register number is 0 the operand is the literal 64-bit zero, not r0. |
RB |
lvsl: read; lvsl128: read | Source GPR. |
VD |
lvsl: write; lvsl128: write | Destination vector register. |
Register Effects
lvsl
- Reads (always):
RA0,RB - Reads (conditional): none
- Writes (always):
VD - Writes (conditional): none
lvsl128
- Reads (always):
RA0,RB - Reads (conditional): none
- Writes (always):
VD - Writes (conditional): none
Status-Register Effects
No condition-register or status-register effects.
Operation (pseudocode)
addr_lo <- ((RA|0) + (RB))[60:63]
for i in 0..15: VD[i] <- addr_lo + i
C Translation Example
/* lvsl VD, RA, RB — load-shift-left permute control */
uint64_t base = (insn.RA == 0) ? 0 : r[insn.RA];
uint8_t sh = (uint8_t)((base + r[insn.RB]) & 0xF);
for (int i = 0; i < 16; ++i) v[insn.VD].b[i] = sh + i;
Implementation References
lvsl
- Canary XML:
tools/ppc-instructions.xml— search formnem="lvsl" - Canary emitter:
src/xenia/cpu/ppc/ppc_emit_altivec.cc:111 - Sylpheed opcode:
crates/sylpheed-ppc/src/opcode.rs:148 - Sylpheed decoder:
crates/sylpheed-ppc/src/decoder.rs:866
Canary emitter (frozen snapshot @ f21ebd49e9)
int InstrEmit_lvsl(PPCHIRBuilder& f, const InstrData& i) {
return InstrEmit_lvsl_(f, i, i.X.RT, i.X.RA, i.X.RB);
}
// ── delegates to (src/xenia/cpu/ppc/ppc_emit_altivec.cc:103) ──
int InstrEmit_lvsl_(PPCHIRBuilder& f, const InstrData& i, uint32_t vd,
uint32_t ra, uint32_t rb) {
Value* ea = CalculateEA_0(f, ra, rb);
Value* sh = f.Truncate(f.And(ea, f.LoadConstantInt64(0xF)), INT8_TYPE);
Value* v = f.LoadVectorShl(sh);
f.StoreVR(vd, v);
return 0;
}
lvsl128
- Canary XML:
tools/ppc-instructions.xml— search formnem="lvsl128" - Canary emitter:
src/xenia/cpu/ppc/ppc_emit_altivec.cc:114 - Sylpheed opcode:
crates/sylpheed-ppc/src/opcode.rs:149 - Sylpheed decoder:
crates/sylpheed-ppc/src/decoder.rs:527
Canary emitter (frozen snapshot @ f21ebd49e9)
int InstrEmit_lvsl128(PPCHIRBuilder& f, const InstrData& i) {
return InstrEmit_lvsl_(f, i, VX128_1_VD128, i.VX128_1.RA, i.VX128_1.RB);
}
// ── delegates to (src/xenia/cpu/ppc/ppc_emit_altivec.cc:103) ──
int InstrEmit_lvsl_(PPCHIRBuilder& f, const InstrData& i, uint32_t vd,
uint32_t ra, uint32_t rb) {
Value* ea = CalculateEA_0(f, ra, rb);
Value* sh = f.Truncate(f.And(ea, f.LoadConstantInt64(0xF)), INT8_TYPE);
Value* v = f.LoadVectorShl(sh);
f.StoreVR(vd, v);
return 0;
}
Extended Pseudocode
; lvsl VD, RA, RB — load vector for shift left (generates a permute mask)
EA <- (RA|0) + (RB) ; full 64-bit EA; only the low 4 bits matter
sh <- EA[60:63] ; bits 60..63 of EA (the misalignment)
for i in 0..15:
VD[i] <- sh + i ; bytes 0..15 of VD = {sh, sh+1, …, sh+15}
Special Cases & Edge Conditions
- No memory is actually read. Despite the name,
lvsl/lvsrdo not touch memory. They consume the effective address only to extract the low four bits (the alignment offset) and materialise a 16-byte permute control vector inVD. They are pure "address → permute-mask" converters. - Big-endian byte indexing.
VD[0]is the most-significant byte of the 128-bit register (lane 0). WhenEA & 0xF == 0the output is{0, 1, 2, …, 15}, i.e. the identity permute. WhenEA & 0xF == 3the output is{3, 4, …, 18}— modulo nothing, the values do exceed 15. That's intentional: fed intovperm(vperm VD, VA, VB, VC), byte selectors 0..15 index intoVAand 16..31 index intoVB. A stream oflvsl+ two alignedlvxloads of consecutive 16-byte blocks +vpermreconstructs the unaligned 16-byte vector atEA. - Pair with
lvsrfor the opposite direction.lvslshifts "left" (toward the low index / high address byte);lvsrshifts "right". Which one to pick depends on which aligned block you're starting from — see the idiom below. - Standard unaligned-load idiom.
lvx vAL, r0, rA ; aligned block at EA & ~0xF lvx vAH, r0, rA + 16 ; next aligned block lvsl vC, r0, rA ; permute mask from misalignment vperm vD, vAL, vAH, vC ; the unaligned 16 bytes starting at EA RA0semantics. WhenRA = 0the base is the literal zero, solvsl vD, 0, rBderives the mask fromrB & 0xF.- VMX128 sibling (
lvsl128). Same semantics; only theVDregister is encoded with the 7-bit VMX128 register-fusion (VD128l ‖ VD128h) sovDmay bev0..v127. - No flags, no side effects beyond writing
VD. Trivial to move and schedule.
Related Instructions
lvsr— the mirror:VD[i] = 16 − sh + i.vperm— consumes the mask to perform arbitrary byte-level permutation across two vectors.lvx,lvlx,lvrx— the actual memory loads used alongside the mask.vsldoi— static-offset shift-double; when the shift is compile-time known, this is cheaper than thelvsl/vpermpair.