[GPU] 8-bit PWL gamma RT as linear 16-bit UNorm on the host
With render target HLE, directly store linear values as R16G16B16A16_UNORM without gamma conversion, as this format provides more than enough bits (need at least 11 per component due to the maximum scale being 2^3 in the piecewise linear gamma curve) to represent linear values without precision loss. This makes blending work correctly in linear space, improving quality of transparency, lighting passes, and fixing issues such as transparent parts of impact and footstep decals in 4D5307E6 being bright instead. The new behavior is enabled by default, as it hugely improves the accuracy of emulation of this format, that is pretty commonplace in Xbox 360 games, with likely just a small GPU memory and bandwidth usage increase, compared to the alternatives that were previously available on the HLE RB path. It's currently implemented only on Direct3D 12, as most of the current GPU emulation code is planned to be phased out and redone, and no methods other than 8-bit with pre-conversion were implemented on Vulkan previously. To implement on Vulkan later, same conversion as in the Direct3D 12 implementation will need to be done in ownership transfer and resolve shaders. Currently it's somewhat inconvenient to decouple the conversion functions in `SpirvShaderTranslator` from an instance of the translator due to vector constant usage. Later, simpler SPIR-V generation functions may be added (`spv::Builder` usage in general is overly verbose). The previously default method (8-bit storage with pre-conversion in shaders and incorrect blending) can be re-enabled by setting the "gamma_render_target_as_unorm16" configuration option to `false`. This may be useful if the game, for instance, switches between 8_8_8_8_GAMMA and 8_8_8_8 formats for the same data frequently, as switching will result in EDRAM range ownership transfer data copying now. Also, the old path is preserved for Vulkan devices not supporting R16G16B16A16_UNORM with blending. The other workaround that was available previously, replacing the PWL encoding with host hardware sRGB with linear-space blending in render target management and in texture fetching, was also inherently inaccurate in many ways (especially when games have their own PWL encoding math, like 4541080F that displayed incorrect colors on the loading screen), and required tracking of the encoding needed for ranges in the memory. The sRGB workaround therefore was deleted in this commit, greatly simplifying the code in the parts of render target, texture and memory management and shader generation that were involved in it.
This commit is contained in:
@@ -95,8 +95,8 @@ class D3D12RenderTargetCache final : public RenderTargetCache {
|
||||
|
||||
// For host render targets.
|
||||
|
||||
bool gamma_render_target_as_srgb() const {
|
||||
return gamma_render_target_as_srgb_;
|
||||
bool gamma_render_target_as_unorm16() const {
|
||||
return gamma_render_target_as_unorm16_;
|
||||
}
|
||||
|
||||
// Using R16G16[B16A16]_SNORM, which are -1...1, not the needed -32...32.
|
||||
@@ -129,6 +129,8 @@ class D3D12RenderTargetCache final : public RenderTargetCache {
|
||||
xenos::DepthRenderTargetFormat format);
|
||||
|
||||
protected:
|
||||
bool IsGammaFormatHostStorageSeparate() const override;
|
||||
|
||||
uint32_t GetMaxRenderTargetWidth() const override {
|
||||
return D3D12_REQ_TEXTURE2D_U_OR_V_DIMENSION;
|
||||
}
|
||||
@@ -225,16 +227,13 @@ class D3D12RenderTargetCache final : public RenderTargetCache {
|
||||
|
||||
class D3D12RenderTarget final : public RenderTarget {
|
||||
public:
|
||||
// descriptor_draw_srgb is only used for k_8_8_8_8 render targets when host
|
||||
// sRGB (gamma_render_target_as_srgb) is used. descriptor_load is present
|
||||
// when the DXGI formats are different for drawing and bit-exact loading
|
||||
// (for NaN pattern preservation across EDRAM tile ownership transfers in
|
||||
// floating-point formats, and to distinguish between two -1 representations
|
||||
// in snorm formats).
|
||||
// descriptor_load is present when the DXGI formats are different for
|
||||
// drawing and bit-exact loading (for NaN pattern preservation across EDRAM
|
||||
// tile ownership transfers in floating-point formats, and to distinguish
|
||||
// between two -1 representations in snorm formats).
|
||||
D3D12RenderTarget(
|
||||
RenderTargetKey key, ID3D12Resource* resource,
|
||||
ui::d3d12::D3D12CpuDescriptorPool::Descriptor&& descriptor_draw,
|
||||
ui::d3d12::D3D12CpuDescriptorPool::Descriptor&& descriptor_draw_srgb,
|
||||
ui::d3d12::D3D12CpuDescriptorPool::Descriptor&&
|
||||
descriptor_load_separate,
|
||||
ui::d3d12::D3D12CpuDescriptorPool::Descriptor&& descriptor_srv,
|
||||
@@ -243,7 +242,6 @@ class D3D12RenderTargetCache final : public RenderTargetCache {
|
||||
: RenderTarget(key),
|
||||
resource_(resource),
|
||||
descriptor_draw_(std::move(descriptor_draw)),
|
||||
descriptor_draw_srgb_(std::move(descriptor_draw_srgb)),
|
||||
descriptor_load_separate_(std::move(descriptor_load_separate)),
|
||||
descriptor_srv_(std::move(descriptor_srv)),
|
||||
descriptor_srv_stencil_(std::move(descriptor_srv_stencil)),
|
||||
@@ -254,10 +252,6 @@ class D3D12RenderTargetCache final : public RenderTargetCache {
|
||||
const {
|
||||
return descriptor_draw_;
|
||||
}
|
||||
const ui::d3d12::D3D12CpuDescriptorPool::Descriptor& descriptor_draw_srgb()
|
||||
const {
|
||||
return descriptor_draw_srgb_;
|
||||
}
|
||||
const ui::d3d12::D3D12CpuDescriptorPool::Descriptor& descriptor_srv()
|
||||
const {
|
||||
return descriptor_srv_;
|
||||
@@ -297,7 +291,6 @@ class D3D12RenderTargetCache final : public RenderTargetCache {
|
||||
private:
|
||||
Microsoft::WRL::ComPtr<ID3D12Resource> resource_;
|
||||
ui::d3d12::D3D12CpuDescriptorPool::Descriptor descriptor_draw_;
|
||||
ui::d3d12::D3D12CpuDescriptorPool::Descriptor descriptor_draw_srgb_;
|
||||
ui::d3d12::D3D12CpuDescriptorPool::Descriptor descriptor_load_separate_;
|
||||
// Texture SRV non-shader-visible descriptors, to prepare shader-visible
|
||||
// descriptors faster, by copying rather than by creating every time.
|
||||
@@ -718,7 +711,7 @@ class D3D12RenderTargetCache final : public RenderTargetCache {
|
||||
|
||||
bool use_stencil_reference_output_ = false;
|
||||
|
||||
bool gamma_render_target_as_srgb_ = false;
|
||||
bool gamma_render_target_as_unorm16_ = false;
|
||||
|
||||
bool depth_float24_round_ = false;
|
||||
bool depth_float24_convert_in_pixel_shader_ = false;
|
||||
@@ -753,7 +746,6 @@ class D3D12RenderTargetCache final : public RenderTargetCache {
|
||||
|
||||
const RenderTarget* const*
|
||||
current_command_list_render_targets_[1 + xenos::kMaxColorRenderTargets];
|
||||
uint32_t are_current_command_list_render_targets_srgb_ = 0;
|
||||
bool are_current_command_list_render_targets_valid_ = false;
|
||||
|
||||
// Temporary storage for descriptors used in PerformTransfersAndResolveClears
|
||||
|
||||
Reference in New Issue
Block a user