Commit Graph

407 Commits

Author SHA1 Message Date
Herman S.
f2b9b57e18 Reapply "[x64] Zero extend on mov to 8bit register"
This reverts commit 3a6f63f34f.

(The issue wasn't xbyak but rather cherrypick order wrt to
other commits, xbyak upgrade was a red herring)
2026-02-14 02:23:58 +09:00
Herman S.
f5afafaec0 Ensure stack allocations maintain 16-byte alignment for AVX instructions 2026-02-14 02:02:40 +09:00
Herman S.
7e66a85f43 [x64/Linux] More return __m128i by value in xmm0 rather than pointer 2026-02-14 01:38:53 +09:00
Herman S.
0921a4fb04 [x64/Linux] Return __m128i by value in xmm0 rather than a pointer 2026-02-14 01:35:12 +09:00
Herman S.
2abf91603e [x64/Linux] Ensure EmitHostToGuestThunk saves rsi 2026-02-14 01:34:59 +09:00
Herman S.
11815400cd [CPU] Fix f16 pack rounding and SHORT_2 test input
Add round-to-nearest-even to the fast float16 pack path by folding a
0xFFF rounding bias into XMMF16PackLCPI0 and extracting bit 13 of the
source as the tie-breaker, matching the software fallback behavior.

Fix PACK_SHORT_2 test to use pre-biased float input (0x40400000 = 3.0)
as the hardware expects, rather than raw 0.0f which is out of range.
2026-02-14 00:48:16 +09:00
Herman S.
bb70e5c651 [CPU] Fix incorrect all_same detection in constant vector shift paths
The loop conditions `n < 8 - n` and `n < 4 - n` terminated early,
only checking the first half of elements. This caused EmitInt16 and
EmitInt32 to incorrectly take the uniform shift path when trailing
elements had different shift amounts.

Resolves potential issues in SHL, SHR, and SHA.
2026-02-14 00:35:09 +09:00
Herman S.
10e8224a63 [CPU] Use RAII lock_guard in GuestTrampolineGroup 2026-02-14 00:17:45 +09:00
Herman S.
dad5f327bf [CPU] Fix off by one in bsearch 2026-02-13 23:33:40 +09:00
Herman S.
3a6f63f34f Revert "[x64] Zero extend on mov to 8bit register"
This reverts commit 88b0aea272.

(It needs to be applied after xbyak update)
2026-02-13 23:13:07 +09:00
Herman S.
1182f9d73d [x64] Add software fallback for PACK_FLOAT16_4
Current implementation has an off by 1 in rounding, should
round up to even but doesn't. Need to figure out how to implement
it properly so just leaving the software version here for later
verification.
2026-02-13 21:48:42 +09:00
Herman S.
374ec3634e [x64] Implement f32 arithmetic and tests
Probably not needed but doesn't hurt to be complete
2026-02-13 21:48:01 +09:00
Herman S.
fdb32b909f [x64] Remove AVX512 optimization for vrefp
The precision is too low so it's more trouble than it's worth.
2026-02-13 21:13:39 +09:00
Herman S.
77e320f79a [x64] Fix AVX2 optimization path for VECTOR_SHL_V128
Ensure the optimized path matches the fallback behavior
2026-02-13 20:40:48 +09:00
Herman S.
b3a131698c [x64] Fix AVX512 optimization path for vadduws 2026-02-13 19:57:44 +09:00
Herman S.
70c1092195 [x64] Fix AVX-512 vctuxs NaN handling
(using ordered comparison predicate)
2026-02-13 19:31:08 +09:00
Herman S.
f95ebb9c55 [x64] Fix mismatched operand sizes in CNTLZ fallback 2026-02-13 19:29:21 +09:00
Herman S.
88b0aea272 [x64] Zero extend on mov to 8bit register 2026-02-13 17:20:31 +09:00
Herman S.
15f61a1a10 [x64] Fix vector mask issue and add missing tests 2026-02-13 17:07:03 +09:00
Herman S.
09a15fbedc [Testing] Get tests running (dirty hack for linux) 2026-02-13 13:32:39 +09:00
Gliniak
14c2814654 [CPU] Fixed bug in VECTOR_SHL_V128 implementation
- For EmitInt8 there was missing check for & 7 which was causing graphical glitches in movies

- There is probably similar bug in 16/32 version, but that's for another commit
2026-02-11 22:41:36 +01:00
Gliniak
702fbc8adb [CPU/X64] Fixed random on boot crashes related to rsp not being aligned
vmovaps returns failure if it encounters unaligned memory.
Alignemnt must be done on 16 bytes, but sometimes rsp ends with 8
2025-11-09 21:59:46 +01:00
Gliniak
f1d444dd99 [CPU/X64] Added const path for indirect call
This prevents having false-positive logs about invalid handling

This case should only happen on deadend paths of code like calling 0
2025-11-01 19:18:45 +01:00
Gliniak
1ad35e4d61 [CPU/X64] Added missing const path to Emit8_IN_16 2025-11-01 17:39:10 +01:00
Gliniak
906c0b8fa6 [CPU] Fixed issue with const path in AVX512 vector left rotate.
- Added logging in case of similar issues
2025-10-28 17:43:21 +01:00
Margen67
25d6f8269a [premake] Cleanup 2025-07-27 02:46:43 -07:00
Margen67
3eef564ff8 [format] Require EOF newline 2025-07-19 14:42:21 -07:00
Margen67
4cb783bf22 Header cleanup 2025-07-19 14:42:21 -07:00
Xphalnos
47f327e848 [Misc] Replaced const with constexpr where possible 2025-04-08 19:32:17 +02:00
Xphalnos
7479ccc292 [Misc] Fix Some Warnings on Clang Build with Windows + Adding constexpr 2025-03-27 17:52:18 +01:00
Xphalnos
5f918ef28d [Misc] Replaced const with constexpr where possible 2025-03-25 19:50:37 +01:00
nikolay-kyosev
9a0ed48168 A fix for the release build crash on linux. 2025-01-27 18:21:14 +01:00
Marco Rodolfi
8d841693ff [cpu] Fix System-V ABI guest to host and host to guest thunk emitters for Linux
Upstream changes made from xenia-project#1339 and xenia-project#2228 back to canary builds. This fixes various emulation crashes caused from different calling conventions on System-V ABI platforms compared to Windows standard.
2025-01-19 16:36:52 +01:00
The-Little-Wolf
b9be601fad [CPU/x64_sequences] - MAX_V128 fixs
- change e.vandps to e.vorps in MAX_V128 to ensure NaN instructions matches real hardware
2025-01-18 20:00:56 +01:00
Gliniak
3ac98ebfba [Static Analysis] Resolved some issues mentioned in Alexandr-u report 2024-11-02 19:08:34 +01:00
Gliniak
ec267c348a [LINT] Fixed lint issues after clang-format update 2024-06-13 20:56:56 +02:00
Gliniak
2ca752ce07 Revert "Optimized CONVERT_I64_TO_F64 with neat overflow trick"
This reverts commit 3ad80810b5.

Fixes broken statue animation is 454108CF
2024-05-09 20:00:03 +02:00
Gliniak
b9061e6292 [LINT] Linted files + Added lint job to CI 2024-03-12 19:19:30 +01:00
Gliniak
354d42135d [CPU] Support for const value in src1 for OPCODE_PACK 2024-02-12 18:33:03 +01:00
Gliniak
c8ca077ec4 [Codacy] Fixed some issues found by codacy.
Added skipping shaders directory in scan
2024-01-20 13:19:37 +01:00
disjtqz
43fd396db7 implement dynamically allocateable guest to host callbacks 2023-10-11 18:37:23 +02:00
disjtqz
fe7dc26e3f place locals on backend pages 2023-10-02 09:50:13 +02:00
disjtqz
67f16c4e31 implement more accurately inaccurate frsqrte 2023-10-02 09:47:30 +02:00
disjtqz
79465708aa implement bit-perfect vrsqrtefp 2023-10-01 11:08:17 +02:00
Gliniak
ce9a82ccf8 Merge branch 'master' of https://github.com/xenia-project/xenia into canary_experimental 2023-09-01 18:20:29 +02:00
Gliniak
c5e6352c34 [CPU] Added constant propagation pass for: OPCODE_AND_NOT 2023-07-27 23:41:45 +03:00
Wunkolo
6ee2e3718f [x64] Add AVX512 optimizations for OPCODE_VECTOR_COMPARE_UGT(Integer)
AVX512 has native unsigned integer comparisons instructions, removing
the need to XOR the most-significant-bit with a constant in memory to
use the signed comparison instructions. These instructions only write to
a k-mask register though and need an additional call to `vpmovm2*` to
turn the mask-register into a vector-mask register.

As of Icelake:
`vpcmpu*` is all L3/T1
`vpmovm2d` is L1/T0.33
`vpmovm2{b,w}` is L3/T0.33

As of Zen4:
`vpcmpu*` is all L3/T0.50
`vpmovm2*` is all L1/T0.25
2023-05-29 14:57:09 -05:00
chss95cs@gmail.com
27c4cef1b5 Added logger flags, for selectively disabling categories of logging (cpu, apu, kernel). Need to make more log messages make use of these flags.
The "close window" keyboard hotkey (Guide-B) now toggles between loglevel -1 and the loglevel set in your config.
Added LoggerBatch class, which accumulates strings into the threads scratch buffer. This is only intended to be used for very high frequency debug logging. if it exhausts the thread buffer, it just silently stops.
Cleaned nearly 8 years of dust off of the pm4 packet disassembler code, now supports all packets that the command processor supports.
Added extremely verbose logging for gpu register writes. This is not compiled in outside of debug builds, requires LogLevel::Debug and log_guest_driven_gpu_register_written_values = true.
Added full logging of all PM4 packets in the cp. This is not compiled in outside of debug builds, requires LogLevel::Debug and disassemble_pm4.
Piggybacked an implementation of guest callstack backtraces using the stackpoints from enable_host_guest_stack_synchronization. If enable_host_guest_stack_synchronization = false, no backtraces can be obtained.
Added log_ringbuffer_kickoff_initiator_bts. when a thread updates the cp's read pointer, it dumps the backtrace of that thread
Changed the names of the gpu registers CALLBACK_ADDRESS and CALLBACK_CONTEXT to the correct names.
Added a note about CP_PROG_COUNTER
Added CP_RB_WPTR to the gpu register table
Added notes about CP_RB_CNTL and CP_RB_RPTR_ADDR. Both aren't necessary for HLE
Changed name of UNKNOWN_0E00 gpu register to TC_CNTL_STATUS. Games only seem to write 1 to it (L2 invalidate)
2023-04-16 12:42:42 -04:00
chss95cs@gmail.com
e75e0425e0 forward branch for double-clear condition in reserved store 2023-04-15 16:22:37 -04:00
chss95cs@gmail.com
7fb4b4cd41 Attempt to emulate reserved load/store more closely. can't do anything for stores of the same value that are done via a non-reserved store to a reserved location
uses a bitmap that splits up the memory space into 65k blocks per bit.  Currently is using the guest virtual address but should be using physical addresses instead.

Currently if a guest does a reserve on a location and then a reserved store to a totally different location we trigger a breakpoint. This should never happen
Also removed the NEGATED_MUL_blah operations. They weren't necessary, nothing special is needed for the negated result variants.

Added a log message for when watched physical memory has a race, it just would be nice to know when it happens and in what games.
2023-04-15 16:06:07 -04:00