[GPU/D3D12] Memexport from anywhere in control flow + 8/16bpp memexport
There's no limit on the number of memory exports in a shader on the real Xenos, and exports can be done anywhere, including in loops. Now, instead of deferring the exports to the end of the shader, and assuming that export allocs are executed only once, Xenia flushes exports when it reaches an alloc (allocs terminate memory exports on Xenos, as well as individual ALU instructions with `serialize`, but not handling this case for simplicity, it's only truly mandatory to flush memory exports before starting a new one), the end of the shader, or a pixel with outstanding exports is killed. To know which eM# registers need to be flushed to the memory, traversing the successors of each exec potentially writing any eM#, and specifying that certain eM# registers might have potentially been written before each reached control flow instruction, until a flush point or the end of the shader is reached. Also, some games export to sub-32bpp formats. These are now supported via atomic AND clearing the bits of the dword to replace followed by an atomic OR inserting the new byte/short.
This commit is contained in:
@@ -134,7 +134,7 @@ class SpirvShaderTranslator : public ShaderTranslator {
|
||||
// (32-bit only - 16-bit indices are always fetched via the Vulkan index
|
||||
// buffer).
|
||||
kSysFlag_VertexIndexLoad = 1u << kSysFlag_VertexIndexLoad_Shift,
|
||||
// For HostVertexShaderTypes kMemexportCompute, kPointListAsTriangleStrip,
|
||||
// For HostVertexShaderTypes kMemExportCompute, kPointListAsTriangleStrip,
|
||||
// kRectangleListAsTriangleStrip, whether the vertex index needs to be
|
||||
// loaded from the index buffer (rather than using autogenerated indices),
|
||||
// and whether it's 32-bit. This is separate from kSysFlag_VertexIndexLoad
|
||||
@@ -427,7 +427,9 @@ class SpirvShaderTranslator : public ShaderTranslator {
|
||||
const ParsedVertexFetchInstruction& instr) override;
|
||||
void ProcessTextureFetchInstruction(
|
||||
const ParsedTextureFetchInstruction& instr) override;
|
||||
void ProcessAluInstruction(const ParsedAluInstruction& instr) override;
|
||||
void ProcessAluInstruction(
|
||||
const ParsedAluInstruction& instr,
|
||||
uint8_t memexport_eM_potentially_written_before) override;
|
||||
|
||||
private:
|
||||
struct TextureBinding {
|
||||
@@ -620,7 +622,7 @@ class SpirvShaderTranslator : public ShaderTranslator {
|
||||
assert_true(edram_fragment_shader_interlock_);
|
||||
return !is_depth_only_fragment_shader_ &&
|
||||
!current_shader().writes_depth() &&
|
||||
!current_shader().is_valid_memexport_used();
|
||||
!current_shader().memexport_eM_written();
|
||||
}
|
||||
void FSI_LoadSampleMask(spv::Id msaa_samples);
|
||||
void FSI_LoadEdramOffsets(spv::Id msaa_samples);
|
||||
|
||||
Reference in New Issue
Block a user