feat(gpu): GPU compute detile for guest tiled textures (Vulkan + Metal) (#592)

* feat(gpu): GPU compute detile for guest tiled textures (Vulkan + Metal)

Move RDNA2 exact-XOR deswizzle (swizzle modes 5/9/24/27, 4bpp) off the CPU
onto a GPU compute pass. GnmTiling.GetDetileParams resolves the shared
addressing into DetileParams; the CPU fallback and both GPU kernels consume
the same params so they never disagree.

Vulkan (verified bit-exact on NVIDIA): SpirvFixedShaders.CreateDetileCompute
hand-emits the SPIR-V kernel; VulkanDetilePass.RecordDetile records the
dispatch into the async batch command buffer (never a blocking submit on the
render thread) with transients retired via fence; VulkanDetileSelfTest
(SHARPEMU_DETILE_SELFTEST=1) checks both entry points against the CPU detile.

Metal (Mac-untested): detile_compute.msl (detile_cs) + MetalDetilePass mirror
the Vulkan pass. The active Metal path CPU-detiles via the new
GnmTiling.DetileWithParams when a texture arrives packaged (empty RgbaPixels +
TiledSource/Detile), keeping Metal correct under default-on with no regression;
wiring MetalDetilePass live is the remaining on-device step.

Flags: GPU detile is default-on (SHARPEMU_GPU_DETILE=0 disables);
[GPU-DETILE] diagnostics gated behind SHARPEMU_LOG_GPU_DETILE=1.

Tests: 17 detile unit tests pass, incl. DetileWithParams and GetDetileParams
each matching TryDetile bit-for-bit across all supported modes/bpp, plus a
SPIR-V structural-validity test.

* feat(gpu): GPU compute detile for guest tiled textures (Vulkan + Metal)

Move RDNA2 exact-XOR deswizzle (swizzle modes 5/9/24/27, 4bpp) off the CPU
onto a GPU compute pass. GnmTiling.GetDetileParams resolves the shared
addressing into DetileParams that the CPU fallback and both GPU kernels
consume, so they never disagree; everything else keeps the CPU path.

Vulkan (verified bit-exact on NVIDIA): SpirvFixedShaders.CreateDetileCompute
hand-emits the kernel; VulkanDetilePass.RecordDetile records into the async
batch command buffer (never a blocking submit on the render thread) with
transients retired via fence, falling back to CPU detile on failure.
VulkanDetileSelfTest (SHARPEMU_DETILE_SELFTEST=1) checks both entry points.

Metal (Mac-untested): detile_compute.msl + MetalDetilePass mirror the Vulkan
pass; the active Metal path CPU-detiles via GnmTiling.DetileWithParams so it
stays correct under default-on. Wiring MetalDetilePass live is a follow-up.

Flags: default-on (SHARPEMU_GPU_DETILE=0 disables); diagnostics behind
SHARPEMU_LOG_GPU_DETILE=1. Adds 17 passing detile unit tests.

* Fix: added support layered texture support for the GPU-Detiling.

* Fix: Added support for BlockTable (1 / 4 / 8 (Morton/Z-order))

* feat: added support for 8 and 16 bpp (bytes per element)

* Fixed a build failure specific to this branch

---------
This commit is contained in:
shadowbeat070
2026-07-24 19:13:02 +02:00
committed by GitHub
parent 5228335f15
commit a158960c20
13 changed files with 2702 additions and 27 deletions
@@ -260,4 +260,233 @@ public static class SpirvFixedShaders
module.AddExecutionMode(main, SpirvExecutionMode.OriginUpperLeft);
return module.Build();
}
/// <summary>
/// Compute kernel that deswizzles RDNA2 tiled surfaces at 4 bytes/element into
/// a linear output buffer — one GPU thread per texel, one dispatch-Z layer per
/// array slice. Mirrors <c>GnmTiling.GetDetileParams</c> so it is bit-identical
/// to the CPU fallback for both supported equation families:
/// <code>
/// z = layer;
/// inBlock = equation == BlockTable // modes 1/4/8
/// ? blockTable[(y % blockHeight) * blockWidth + (x % blockWidth)]
/// : xTerm[x &amp; xMask] ^ yTerm[y &amp; yMask]; // ExactXor 5/9/24/27
/// src = z * srcSliceElements
/// + (y / blockHeight * blocksPerRow + x / blockWidth) * blockElements
/// + inBlock;
/// out[z * width * height + y * width + x] = tiled[src];
/// </code>
/// Each array slice is an independently tiled 2D surface; the caller packs the
/// slices contiguously in the tiled buffer (stride <c>srcSliceElements</c>) and
/// the output ends up layer-major, matching a single multi-layer
/// buffer-&gt;image copy. For a non-arrayed texture the caller dispatches a
/// single Z layer with <c>srcSliceElements</c> unused (z == 0).
///
/// The term tables hold ELEMENT offsets. For ExactXor the caller pre-shifts the
/// byte-unit GetDetileParams terms right by log2(bytesPerElement) (exact at 4bpp
/// since the equation's low two byte-offset bits are 0); for BlockTable the
/// GetDetileParams block table is already in element units. Binding 1 carries
/// xTerm (ExactXor) OR blockTable (BlockTable) — the two equations index
/// different-sized buffers, so the kernel branches and evaluates exactly one.
///
/// width/height are ELEMENT dims (for block-compressed formats a 4x4 block is
/// one element). Each element spans uintsPerElement = bpp/4 words (4bpp -> 1,
/// 8bpp -> 2, 16bpp -> 4); the X dispatch is widened by that factor so each
/// thread copies one word (elemX = gidX / upe, word = gidX % upe). 1/2 bpp are
/// sub-word and stay on the CPU.
///
/// Descriptor set 0: binding 0 = tiled uint[], 1 = xTerm/blockTable uint[],
/// 2 = yTerm uint[], 3 = out uint[]. Push constants (11 x uint, offset i*4):
/// width, height, blockWidth, blockHeight, blockElements, blocksPerRow,
/// xMask, yMask, srcSliceElements, equation (0 = ExactXor, 1 = BlockTable),
/// uintsPerElement. Local size 8x8x1; dispatch X = ceil(width*upe/8),
/// Y = ceil(height/8), Z = arrayLayers.
/// </summary>
public static byte[] CreateDetileCompute()
{
var module = new SpirvModuleBuilder();
module.AddCapability(SpirvCapability.Shader);
var voidType = module.TypeVoid();
var boolType = module.TypeBool();
var uintType = module.TypeInt(32, signed: false);
var uvec3Type = module.TypeVector(uintType, 3);
// One shared Block-decorated storage-buffer struct: struct { uint data[]; }.
var runtimeArray = module.TypeRuntimeArray(uintType);
module.AddDecoration(runtimeArray, SpirvDecoration.ArrayStride, 4);
var bufferStruct = module.TypeStruct(runtimeArray);
module.AddDecoration(bufferStruct, SpirvDecoration.Block);
module.AddMemberDecoration(bufferStruct, 0, SpirvDecoration.Offset, 0);
var bufferPtrType = module.TypePointer(SpirvStorageClass.StorageBuffer, bufferStruct);
var uintStoragePtr = module.TypePointer(SpirvStorageClass.StorageBuffer, uintType);
uint MakeBuffer(uint binding, string name)
{
var variable = module.AddGlobalVariable(bufferPtrType, SpirvStorageClass.StorageBuffer);
module.AddName(variable, name);
module.AddDecoration(variable, SpirvDecoration.DescriptorSet, 0);
module.AddDecoration(variable, SpirvDecoration.Binding, binding);
return variable;
}
var tiledVar = MakeBuffer(0, "tiled");
var xTermVar = MakeBuffer(1, "xTerm");
var yTermVar = MakeBuffer(2, "yTerm");
var outVar = MakeBuffer(3, "outLinear");
// Push constants: struct { uint p0..p10; }, each member at offset i*4.
var pushStruct = module.TypeStruct(
uintType, uintType, uintType, uintType, uintType, uintType,
uintType, uintType, uintType, uintType, uintType);
module.AddDecoration(pushStruct, SpirvDecoration.Block);
for (uint member = 0; member < 11; member++)
{
module.AddMemberDecoration(pushStruct, member, SpirvDecoration.Offset, member * 4);
}
var pushPtrType = module.TypePointer(SpirvStorageClass.PushConstant, pushStruct);
var pushMemberPtrType = module.TypePointer(SpirvStorageClass.PushConstant, uintType);
var pushVar = module.AddGlobalVariable(pushPtrType, SpirvStorageClass.PushConstant);
module.AddName(pushVar, "pc");
var inputUvec3Ptr = module.TypePointer(SpirvStorageClass.Input, uvec3Type);
var gidVar = module.AddGlobalVariable(inputUvec3Ptr, SpirvStorageClass.Input);
module.AddName(gidVar, "gid");
module.AddDecoration(gidVar, SpirvDecoration.BuiltIn, (uint)SpirvBuiltIn.GlobalInvocationId);
var uintConst = new uint[11];
for (uint value = 0; value < 11; value++)
{
uintConst[value] = module.Constant(uintType, value);
}
var functionType = module.TypeFunction(voidType);
var main = module.BeginFunction(voidType, functionType);
module.AddName(main, "main");
module.AddLabel();
var gid = module.AddInstruction(SpirvOp.Load, uvec3Type, gidVar);
var gidX = module.AddInstruction(SpirvOp.CompositeExtract, uintType, gid, 0);
var y = module.AddInstruction(SpirvOp.CompositeExtract, uintType, gid, 1);
var z = module.AddInstruction(SpirvOp.CompositeExtract, uintType, gid, 2);
uint PushField(uint index)
{
var pointer = module.AddInstruction(
SpirvOp.AccessChain, pushMemberPtrType, pushVar, uintConst[index]);
return module.AddInstruction(SpirvOp.Load, uintType, pointer);
}
// width/height are ELEMENT dims (for BC, a 4x4 block is one element). Each
// element spans uintsPerElement 32-bit words (bpp/4: 4bpp->1, 8bpp->2,
// 16bpp->4). The X dispatch is widened by uintsPerElement so each thread
// copies exactly one word: elemX = gidX / upe, wordIndex = gidX % upe.
var width = PushField(0);
var height = PushField(1);
var blockWidth = PushField(2);
var blockHeight = PushField(3);
var blockElements = PushField(4);
var blocksPerRow = PushField(5);
var xMask = PushField(6);
var yMask = PushField(7);
var srcSliceElements = PushField(8);
var equation = PushField(9);
var uintsPerElement = PushField(10);
var elemX = module.AddInstruction(SpirvOp.UDiv, uintType, gidX, uintsPerElement);
var elemXTimesUpe = module.AddInstruction(SpirvOp.IMul, uintType, elemX, uintsPerElement);
var wordIndex = module.AddInstruction(SpirvOp.ISub, uintType, gidX, elemXTimesUpe);
var xInRange = module.AddInstruction(SpirvOp.ULessThan, boolType, elemX, width);
var yInRange = module.AddInstruction(SpirvOp.ULessThan, boolType, y, height);
var inRange = module.AddInstruction(SpirvOp.LogicalAnd, boolType, xInRange, yInRange);
var bodyLabel = module.AllocateId();
var mergeLabel = module.AllocateId();
module.AddStatement(SpirvOp.SelectionMerge, mergeLabel, 0);
module.AddStatement(SpirvOp.BranchConditional, inRange, bodyLabel, mergeLabel);
module.AddLabel(bodyLabel);
// blockIdx = (y / blockHeight) * blocksPerRow + (elemX / blockWidth)
var yDiv = module.AddInstruction(SpirvOp.UDiv, uintType, y, blockHeight);
var blockRow = module.AddInstruction(SpirvOp.IMul, uintType, yDiv, blocksPerRow);
var xDiv = module.AddInstruction(SpirvOp.UDiv, uintType, elemX, blockWidth);
var blockIdx = module.AddInstruction(SpirvOp.IAdd, uintType, blockRow, xDiv);
// off (element offset within the block) = equation == BlockTable
// ? blockTable[(y % blockHeight) * blockWidth + (elemX % blockWidth)]
// : xTerm[elemX & xMask] ^ yTerm[y & yMask]
// Binding 1 (xTermVar) doubles as the block table; the two equations index
// different-sized buffers, so exactly one branch executes (no OOB read).
var isBlockTable = module.AddInstruction(SpirvOp.INotEqual, boolType, equation, uintConst[0]);
var xorLabel = module.AllocateId();
var tableLabel = module.AllocateId();
var offMergeLabel = module.AllocateId();
module.AddStatement(SpirvOp.SelectionMerge, offMergeLabel, 0);
module.AddStatement(SpirvOp.BranchConditional, isBlockTable, tableLabel, xorLabel);
// ExactXor: xTerm[elemX & xMask] ^ yTerm[y & yMask]
module.AddLabel(xorLabel);
var xIdx = module.AddInstruction(SpirvOp.BitwiseAnd, uintType, elemX, xMask);
var xPtr = module.AddInstruction(SpirvOp.AccessChain, uintStoragePtr, xTermVar, uintConst[0], xIdx);
var xTerm = module.AddInstruction(SpirvOp.Load, uintType, xPtr);
var yIdx = module.AddInstruction(SpirvOp.BitwiseAnd, uintType, y, yMask);
var yPtr = module.AddInstruction(SpirvOp.AccessChain, uintStoragePtr, yTermVar, uintConst[0], yIdx);
var yTerm = module.AddInstruction(SpirvOp.Load, uintType, yPtr);
var offXor = module.AddInstruction(SpirvOp.BitwiseXor, uintType, xTerm, yTerm);
module.AddStatement(SpirvOp.Branch, offMergeLabel);
// BlockTable: blockTable[inY * blockWidth + inX], inX/inY = position in block
module.AddLabel(tableLabel);
var blockXBase = module.AddInstruction(SpirvOp.IMul, uintType, xDiv, blockWidth);
var inX = module.AddInstruction(SpirvOp.ISub, uintType, elemX, blockXBase);
var blockYBase = module.AddInstruction(SpirvOp.IMul, uintType, yDiv, blockHeight);
var inY = module.AddInstruction(SpirvOp.ISub, uintType, y, blockYBase);
var rowInBlock = module.AddInstruction(SpirvOp.IMul, uintType, inY, blockWidth);
var tableIdx = module.AddInstruction(SpirvOp.IAdd, uintType, rowInBlock, inX);
var tablePtr = module.AddInstruction(SpirvOp.AccessChain, uintStoragePtr, xTermVar, uintConst[0], tableIdx);
var offTable = module.AddInstruction(SpirvOp.Load, uintType, tablePtr);
module.AddStatement(SpirvOp.Branch, offMergeLabel);
module.AddLabel(offMergeLabel);
var off = module.AddInstruction(SpirvOp.Phi, uintType, offXor, xorLabel, offTable, tableLabel);
// srcElem = z * srcSliceElements + blockIdx * blockElements + off (in elements)
// srcWord = srcElem * uintsPerElement + wordIndex
var srcSliceBase = module.AddInstruction(SpirvOp.IMul, uintType, z, srcSliceElements);
var blockBase = module.AddInstruction(SpirvOp.IMul, uintType, blockIdx, blockElements);
var srcInSlice = module.AddInstruction(SpirvOp.IAdd, uintType, blockBase, off);
var srcElem = module.AddInstruction(SpirvOp.IAdd, uintType, srcSliceBase, srcInSlice);
var srcElemWords = module.AddInstruction(SpirvOp.IMul, uintType, srcElem, uintsPerElement);
var src = module.AddInstruction(SpirvOp.IAdd, uintType, srcElemWords, wordIndex);
var srcPtr = module.AddInstruction(SpirvOp.AccessChain, uintStoragePtr, tiledVar, uintConst[0], src);
var word = module.AddInstruction(SpirvOp.Load, uintType, srcPtr);
// dstElem = z * width * height + y * width + elemX (in elements)
// dstWord = dstElem * uintsPerElement + wordIndex
var sliceElements = module.AddInstruction(SpirvOp.IMul, uintType, width, height);
var dstSliceBase = module.AddInstruction(SpirvOp.IMul, uintType, z, sliceElements);
var rowBase = module.AddInstruction(SpirvOp.IMul, uintType, y, width);
var dstRow = module.AddInstruction(SpirvOp.IAdd, uintType, rowBase, elemX);
var dstElem = module.AddInstruction(SpirvOp.IAdd, uintType, dstSliceBase, dstRow);
var dstElemWords = module.AddInstruction(SpirvOp.IMul, uintType, dstElem, uintsPerElement);
var dstIdx = module.AddInstruction(SpirvOp.IAdd, uintType, dstElemWords, wordIndex);
var dstPtr = module.AddInstruction(SpirvOp.AccessChain, uintStoragePtr, outVar, uintConst[0], dstIdx);
module.AddStatement(SpirvOp.Store, dstPtr, word);
module.AddStatement(SpirvOp.Branch, mergeLabel);
module.AddLabel(mergeLabel);
module.AddStatement(SpirvOp.Return);
module.EndFunction();
module.AddExecutionMode(main, SpirvExecutionMode.LocalSize, 8, 8, 1);
module.AddEntryPoint(
SpirvExecutionModel.GLCompute,
main,
"main",
[gidVar, tiledVar, xTermVar, yTermVar, outVar, pushVar]);
return module.Build();
}
}