[Gpu] Backend-neutral shader compiler and guest-GPU renderer seam (#200)

* [ShaderCompiler] Extract the backend-neutral shader compiler project

Move the Gen5 (gfx10) microcode decoder, the scalar evaluator, the
shader IR, and the metadata reader out of SharpEmu.Libs/Agc into a new
SharpEmu.ShaderCompiler project — the half of shader compilation every
codegen backend (SPIR-V today; MSL and DXIL later) consumes. Types go
public: they are the contract now. Nothing in the project may depend on
a host graphics API; the SPIR-V-specific artifact types
(Gen5SpirvShader, Gen5SpirvStage) stay beside the emitter in Libs.

Three couplings surfaced by the move, each resolved at the right depth:
GuestDrawKind was defined inside VulkanVideoPresenter despite being a
guest-domain, decoder-produced concept — it moves to the shared project;
the evaluator's one HLE dependency (the tracked-libc-heap read
fallback) becomes an injectable hook that a Libs module initializer
installs before any caller can reach the evaluator; and the inline-
constant table is promoted to a shared Gen5InlineConstants so backends
cannot drift on constant semantics (the SPIR-V translator now delegates
to it).

The ShaderDump tool drops its reflection over the moved types in favor
of direct typed calls; only the SPIR-V emitter, still internal to Libs
until it moves to its own backend project, is reached via reflection.
Verified by a clean solution build, the existing test suite, and a full
ShaderDump conformance run.

* [ShaderCompiler] Move the SPIR-V emitter into SharpEmu.ShaderCompiler.Vulkan

Gen5SpirvTranslator (with its ALU partial), SpirvModuleBuilder,
SpirvFixedShaders, and the Gen5SpirvShader/Gen5SpirvStage artifact types
move whole from SharpEmu.Libs/Agc into the first per-backend codegen
project. Notably it needs no Vulkan bindings reference: emitters
produce bytes from the shared IR; renderers own graphics APIs. Types go
public as the backend's contract; AgcExports and the presenter consume
them exactly as before.

The ShaderDump tool drops its last reflection: with both halves of the
pipeline public it drives decode and all three emit entry points with
direct typed calls, retiring the PadWithDefaults invoke shim — and it
no longer references SharpEmu.Libs at all, making the conformance tool
emulator-independent by design. Verified by a clean solution build, the
test suite, a full ShaderDump conformance run, and a locked-mode
restore under the pinned SDK.

* [Gpu] Extract the guest-GPU backend seam (IGuestGpuBackend)

The AGC/VideoOut/SystemService export layers now reach the renderer
through IGuestGpuBackend via GuestGpu.Current (mirroring HostPlatform),
instead of calling VulkanVideoPresenter statics. The Vulkan backend is
a thin adapter over the existing presenter, so the extraction stays
mechanical; only the adapter and the presenter itself reference the
presenter now.

The types crossing the seam move to Gpu/GuestGpuTypes.cs and drop their
Vulkan prefixes, which an audit showed were misnomers: every field is a
neutral primitive or a raw guest value (guest addresses, format and
number-type codes, CB_BLEND register bitfields, verbatim sampler
descriptor dwords). The one genuine Vulkan value in the old surface —
the Silk.NET Format inside VulkanRenderTargetFormat, which callers
never read — stops crossing: TryDecodeRenderTargetFormat is replaced at
the seam by TryGetRenderTargetOutputKind, which surfaces only the
Gen5PixelOutputKind callers actually consume, keeping native formats a
backend-internal concern. ToVulkanSampler in AgcExports is renamed
ToGuestSampler to match what it always produced.

Seam rules are documented on the interface: no host-API value crosses,
and submission stays coarse-grained with synchronization internal to
backends. Interim exception, resolved next: shader parameters are still
SPIR-V blobs.

* [Gpu] Move shader compilation behind the guest-GPU backend

The seam's interim exception is gone: AgcExports no longer calls
Gen5SpirvTranslator or handles SPIR-V bytes. IGuestGpuBackend gains the
three TryCompile entry points, which take the backend-neutral
(Gen5ShaderState, Gen5ShaderEvaluation) contract plus the flat
per-role resource-slot bases a multi-stage draw needs, and return
opaque IGuestCompiledShader handles that only the producing backend can
submit — the Vulkan backend wraps its SPIR-V in
VulkanCompiledGuestShader and rejects foreign handles loudly. Draw and
dispatch submissions take handles instead of byte arrays; the shader
caches in AgcExports store handles.

IGuestCompiledShader.Payload exposes the backend-defined compiled bytes
for exactly two callers: the diagnostics dump and the size trace —
documented as never-interpret. The unused _pixelSpirvCache is deleted.
With this, a Metal or DX12 backend plugs in by implementing
IGuestGpuBackend with its own codegen; nothing in the export layers
knows which shader format exists.

Verified by a clean solution build, the test suite, and a full
ShaderDump conformance run under the pinned SDK.

* [Gpu] Fix rename collateral from the seam extraction

Address review findings: a doc comment picked up the mechanical
VulkanVideoPresenter -> GuestGpu.Current rewrite and ended up naming
members that do not exist on the interface, and CreateVulkanIndexBuffer
kept its Vulkan prefix while every sibling factory was de-Vulkanized —
it produces the neutral GuestIndexBuffer, so it is CreateGuestIndexBuffer.

* [Gpu] Label diagnostics dumps with the backend's payload extension

Address the review's altitude finding on DumpSpirv: the dump helper's
IR-disassembly half is backend-neutral and stays put, but writing the
opaque payload to a hardcoded .spv interpreted bytes the seam says
never to interpret. IGuestCompiledShader now declares its payload's
file extension, and the renamed DumpCompiledShader takes the handle and
writes honestly-labeled dumps whichever backend produced them.

* [Gpu] Make the shader-cache hit path allocation-free and lock-free

Every translated draw built its cache key with a LINQ Select feeding
string.Join plus one interpolated string per render target — steady
per-draw allocation whether or not the shaders were already cached. The
output layout is now packed exactly into a ulong (guest slot in 6 bits
+ output kind in 2 bits per target, host locations being the byte
positions, target count in the key beside it), and the
Gen5PixelOutputBinding array is only materialized on a cache miss,
where compilation dwarfs it.

The graphics/compute shader caches switch from Dictionary guarded by
_submitTraceGate to ConcurrentDictionary, making the per-draw and
per-dispatch hit paths lock-free and decoupling them from the tracing
gate they coincidentally shared. And the seam-shaped render-target list
is built once when a translated draw is created instead of a
Select/ToArray per submission of a cached draw.

* [Gpu] Replace LINQ with explicit loops in code this branch introduced

Project rule going forward: no LINQ — it allocates enumerators,
closures, and delegates, and this codebase is GC-pause-sensitive. The
pixel-output and guest-render-target array builds and the ShaderDump
store-PC collection become plain loops; pre-existing LINQ elsewhere is
left for changes that already touch those lines.

* [ShaderCompiler] Suppress CA2255 on the evaluator hook installer

The analyzer coverage that arrived with the rebase flags
ModuleInitializer in library code; this is the rule's intended advanced
scenario — the hook must be installed before any code path can reach
the evaluator, and every such path enters through this assembly — so
suppress with that justification rather than weaken the guarantee to a
static constructor's lazier timing.

* [Gpu] Resolve rebase artifacts onto main

Dedupe the System.Collections.Concurrent using in AgcExports that the
rebase merge duplicated (main and this branch each added it), and
regenerate the lock files for the new shader-compiler projects and
SharpEmu.Libs against main's current package graph so --locked-mode
restore matches at the branch tip.

* [CI] Comment per-platform build artifact links on PRs

Adds a workflow_run workflow that, after "Build and Release" finishes a
pull-request build, posts (and keeps updated in place) a single PR
comment linking the Windows, Linux, and macOS artifacts from that run.

It runs via workflow_run rather than in the build workflow because PRs
from forks build with a read-only token that cannot comment; the
follow-on run executes in the base-repo context with write access and
without checking out fork code. GitHub only triggers workflow_run from
the default branch, so this takes effect once merged to main.
This commit is contained in:
Gutemberg Ribeiro
2026-07-15 18:11:24 +01:00
committed by GitHub
parent c69ac6ddab
commit 30fdd8d6ed
35 changed files with 1308 additions and 646 deletions
@@ -0,0 +1,54 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
namespace SharpEmu.ShaderCompiler;
/// <summary>
/// The Gen5 (gfx10) inline-constant operand table, shared by every codegen so backends
/// cannot drift on constant semantics.
/// </summary>
public static class Gen5InlineConstants
{
public static bool TryDecode(uint encoded, out uint value)
{
if (encoded == 125)
{
value = 0;
return true;
}
if (encoded is >= 128 and <= 192)
{
value = encoded - 128;
return true;
}
if (encoded is >= 193 and <= 208)
{
value = unchecked((uint)-(int)(encoded - 192));
return true;
}
var floatingPoint = encoded switch
{
240 => 0.5f,
241 => -0.5f,
242 => 1.0f,
243 => -1.0f,
244 => 2.0f,
245 => -2.0f,
246 => 4.0f,
247 => -4.0f,
248 => 1.0f / (2.0f * MathF.PI),
_ => float.NaN,
};
if (float.IsNaN(floatingPoint))
{
value = 0;
return false;
}
value = BitConverter.SingleToUInt32Bits(floatingPoint);
return true;
}
}
+304
View File
@@ -0,0 +1,304 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
namespace SharpEmu.ShaderCompiler;
public enum Gen5ShaderEncoding
{
Sop1,
Sop2,
Sopc,
Sopp,
Sopk,
Smrd,
Smem,
Mubuf,
Mtbuf,
Vop1,
Vop2,
Vopc,
Vop3,
Vintrp,
Ds,
Flat,
Vop3p,
Mimg,
Exp,
}
public enum Gen5OperandKind
{
ScalarRegister,
VectorRegister,
EncodedConstant,
LiteralConstant,
}
public enum Gen5ShaderResourceKind
{
ReadOnlyTexture,
ReadWriteTexture,
Sampler,
ConstantBuffer,
}
public enum Gen5PixelOutputKind
{
Float,
Uint,
Sint,
}
public readonly record struct Gen5PixelOutputBinding(
uint GuestSlot,
uint HostLocation,
Gen5PixelOutputKind Kind);
public readonly record struct Gen5ShaderResourceMapping(
Gen5ShaderResourceKind Kind,
uint Slot,
uint OffsetDwords,
bool SizeFlag);
public sealed record Gen5ShaderMetadata(
uint ExtendedUserDataSizeDwords,
uint ShaderResourceTableSizeDwords,
IReadOnlyDictionary<uint, uint> DirectResources,
IReadOnlyList<Gen5ShaderResourceMapping> Resources);
public readonly record struct Gen5ComputeSystemRegisters(
uint? WorkGroupXRegister,
uint? WorkGroupYRegister,
uint? WorkGroupZRegister,
uint? ThreadGroupSizeRegister)
{
public bool TryGetExpression(uint scalarRegister, out string expression)
{
if (WorkGroupXRegister == scalarRegister)
{
expression = "gl_WorkGroupID.x";
return true;
}
if (WorkGroupYRegister == scalarRegister)
{
expression = "gl_WorkGroupID.y";
return true;
}
if (WorkGroupZRegister == scalarRegister)
{
expression = "gl_WorkGroupID.z";
return true;
}
if (ThreadGroupSizeRegister == scalarRegister)
{
expression = "(gl_WorkGroupSize.x * gl_WorkGroupSize.y * gl_WorkGroupSize.z)";
return true;
}
expression = string.Empty;
return false;
}
public void ClearStaticValues(Span<uint> scalarRegisters)
{
ClearStaticValue(scalarRegisters, WorkGroupXRegister);
ClearStaticValue(scalarRegisters, WorkGroupYRegister);
ClearStaticValue(scalarRegisters, WorkGroupZRegister);
ClearStaticValue(scalarRegisters, ThreadGroupSizeRegister);
}
private static void ClearStaticValue(Span<uint> scalarRegisters, uint? scalarRegister)
{
if (scalarRegister is { } register && register < scalarRegisters.Length)
{
scalarRegisters[(int)register] = 0;
}
}
}
public sealed record Gen5ShaderState(
Gen5ShaderProgram Program,
IReadOnlyList<uint> UserData,
Gen5ShaderMetadata? Metadata,
Gen5ComputeSystemRegisters? ComputeSystemRegisters = null,
uint UserDataScalarRegisterBase = 0);
public readonly record struct Gen5Operand(Gen5OperandKind Kind, uint Value)
{
public static Gen5Operand Scalar(uint index) =>
new(Gen5OperandKind.ScalarRegister, index);
public static Gen5Operand Vector(uint index) =>
new(Gen5OperandKind.VectorRegister, index);
public static Gen5Operand Source(uint encoded, uint? literal = null)
{
if (encoded >= 256)
{
return Vector(encoded - 256);
}
if (encoded is 249 or 255 && literal.HasValue)
{
return new(Gen5OperandKind.LiteralConstant, literal.Value);
}
if (encoded <= 105 || encoded is 106 or 107 or 124 or 126 or 127)
{
return Scalar(encoded);
}
return new(Gen5OperandKind.EncodedConstant, encoded);
}
public override string ToString() => Kind switch
{
Gen5OperandKind.ScalarRegister => $"s{Value}",
Gen5OperandKind.VectorRegister => $"v{Value}",
Gen5OperandKind.LiteralConstant => $"0x{Value:X8}",
_ => $"src[{Value}]",
};
}
public abstract record Gen5InstructionControl;
public sealed record Gen5ImageControl(
uint Dmask,
uint VectorAddress,
IReadOnlyList<uint> AddressRegisters,
uint VectorData,
uint ScalarResource,
uint ScalarSampler,
uint Dimension,
bool IsArray,
bool Glc,
bool Slc) : Gen5InstructionControl
{
public uint GetAddressRegister(int component) =>
component < AddressRegisters.Count
? AddressRegisters[component]
: VectorAddress + (uint)component;
}
public sealed record Gen5GlobalMemoryControl(
uint DwordCount,
uint VectorAddress,
uint VectorData,
uint ScalarAddress,
int OffsetBytes,
bool Glc,
bool Slc) : Gen5InstructionControl;
public sealed record Gen5BufferMemoryControl(
uint DwordCount,
uint VectorAddress,
uint VectorData,
uint ScalarResource,
int OffsetBytes,
bool IndexEnabled,
bool OffsetEnabled,
bool Glc,
bool Slc) : Gen5InstructionControl;
public sealed record Gen5ExportControl(
uint Target,
uint EnableMask,
bool Compressed,
bool Done,
bool ValidMask) : Gen5InstructionControl;
public sealed record Gen5InterpolationControl(
uint Attribute,
uint Channel) : Gen5InstructionControl;
public sealed record Gen5Vop3Control(
uint AbsoluteMask,
uint NegateMask,
uint OutputModifier,
bool Clamp,
uint? ScalarDestination) : Gen5InstructionControl;
public sealed record Gen5SdwaControl(
uint DestinationSelect,
uint Source0Select,
uint Source1Select,
uint AbsoluteMask,
uint NegateMask,
uint OutputModifier,
bool Clamp) : Gen5InstructionControl;
public sealed record Gen5DppControl(
uint Control,
bool FetchInactive,
bool BoundControl,
uint AbsoluteMask,
uint NegateMask,
uint BankMask,
uint RowMask) : Gen5InstructionControl;
public sealed record Gen5ScalarMemoryControl(
uint DestinationCount,
int ImmediateOffsetBytes,
uint? DynamicOffsetRegister) : Gen5InstructionControl;
public sealed record Gen5DataShareControl(
uint Offset0,
uint Offset1,
bool Gds) : Gen5InstructionControl;
public sealed record Gen5ImageBinding(
uint Pc,
string Opcode,
Gen5ImageControl Control,
IReadOnlyList<uint> ResourceDescriptor,
IReadOnlyList<uint> SamplerDescriptor,
uint? MipLevel);
public sealed record Gen5GlobalMemoryBinding(
uint ScalarAddress,
ulong BaseAddress,
IReadOnlyList<uint> InstructionPcs,
byte[] Data);
public sealed record Gen5VertexInputBinding(
uint Pc,
uint Location,
uint ComponentCount,
uint DataFormat,
uint NumberFormat,
ulong BaseAddress,
uint Stride,
uint OffsetBytes,
byte[] Data);
public sealed record Gen5ShaderEvaluation(
IReadOnlyList<uint> InitialScalarRegisters,
IReadOnlyList<uint> ScalarRegisters,
IReadOnlyDictionary<uint, IReadOnlyList<uint>> ScalarRegistersByPc,
IReadOnlyList<Gen5ImageBinding> ImageBindings,
IReadOnlyList<Gen5GlobalMemoryBinding> GlobalMemoryBindings,
Gen5ComputeSystemRegisters? ComputeSystemRegisters = null,
IReadOnlySet<uint>? RuntimeScalarRegisters = null,
IReadOnlyList<Gen5VertexInputBinding>? VertexInputs = null);
public sealed record Gen5ShaderInstruction(
uint Pc,
Gen5ShaderEncoding Encoding,
string Opcode,
IReadOnlyList<uint> Words,
IReadOnlyList<Gen5Operand> Sources,
IReadOnlyList<Gen5Operand> Destinations,
Gen5InstructionControl? Control);
public sealed record Gen5ShaderProgram(
ulong Address,
IReadOnlyList<Gen5ShaderInstruction> Instructions)
{
public IEnumerable<Gen5ImageControl> ImageResources =>
Instructions
.Select(instruction => instruction.Control)
.OfType<Gen5ImageControl>();
}
@@ -0,0 +1,124 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using SharpEmu.HLE;
namespace SharpEmu.ShaderCompiler;
public static class Gen5ShaderMetadataReader
{
private const ulong ShaderUserDataOffset = 0x08;
private const int ResourceClassCount = 4;
private const int MaxMetadataEntries = 4096;
public static bool TryRead(
CpuContext ctx,
ulong shaderHeaderAddress,
out Gen5ShaderMetadata metadata)
{
metadata = default!;
if (!ctx.TryReadUInt64(shaderHeaderAddress + ShaderUserDataOffset, out var userDataAddress) ||
userDataAddress == 0 ||
!ctx.TryReadUInt64(userDataAddress, out var directResourceOffsetsAddress))
{
return false;
}
var resourceOffsets = new ulong[ResourceClassCount];
for (var resourceClass = 0; resourceClass < ResourceClassCount; resourceClass++)
{
if (!ctx.TryReadUInt64(
userDataAddress + 0x08 + (ulong)(resourceClass * sizeof(ulong)),
out resourceOffsets[resourceClass]))
{
return false;
}
}
if (!ctx.TryReadUInt16(userDataAddress + 0x28, out var extendedUserDataSize) ||
!ctx.TryReadUInt16(userDataAddress + 0x2A, out var shaderResourceTableSize) ||
!ctx.TryReadUInt16(userDataAddress + 0x2C, out var directResourceCount) ||
directResourceCount > MaxMetadataEntries)
{
return false;
}
var resourceCounts = new ushort[ResourceClassCount];
for (var resourceClass = 0; resourceClass < ResourceClassCount; resourceClass++)
{
if (!ctx.TryReadUInt16(
userDataAddress + 0x2E + (ulong)(resourceClass * sizeof(ushort)),
out resourceCounts[resourceClass]) ||
resourceCounts[resourceClass] > MaxMetadataEntries)
{
return false;
}
}
var directResources = new Dictionary<uint, uint>();
if (directResourceCount != 0)
{
if (directResourceOffsetsAddress == 0)
{
return false;
}
for (uint type = 0; type < directResourceCount; type++)
{
if (!ctx.TryReadUInt16(directResourceOffsetsAddress + type * sizeof(ushort), out var offset))
{
return false;
}
if (offset != ushort.MaxValue)
{
directResources[type] = offset;
}
}
}
var resources = new List<Gen5ShaderResourceMapping>();
for (var resourceClass = 0; resourceClass < ResourceClassCount; resourceClass++)
{
var count = resourceCounts[resourceClass];
if (count == 0)
{
continue;
}
if (resourceOffsets[resourceClass] == 0)
{
return false;
}
for (uint slot = 0; slot < count; slot++)
{
if (!ctx.TryReadUInt16(
resourceOffsets[resourceClass] + slot * sizeof(ushort),
out var sharp))
{
return false;
}
var offset = (uint)(sharp & 0x7FFF);
if (offset == 0x7FFF)
{
continue;
}
resources.Add(new Gen5ShaderResourceMapping(
(Gen5ShaderResourceKind)resourceClass,
slot,
offset,
(sharp & 0x8000) != 0));
}
}
metadata = new Gen5ShaderMetadata(
extendedUserDataSize,
shaderResourceTableSize,
directResources,
resources);
return true;
}
}
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,15 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
namespace SharpEmu.ShaderCompiler;
/// <summary>
/// Guest draw patterns the decoder recognizes from known shader programs. Guest-domain,
/// backend-neutral — it previously lived inside the Vulkan presenter, which is exactly
/// the kind of placement this project exists to prevent.
/// </summary>
public enum GuestDrawKind
{
None,
FullscreenBarycentric,
}
@@ -0,0 +1,20 @@
<!--
Copyright (C) 2026 SharpEmu Emulator Project
SPDX-License-Identifier: GPL-2.0-or-later
-->
<Project Sdk="Microsoft.NET.Sdk">
<!-- The backend-neutral half of guest shader compilation: the Gen5 (gfx10) microcode
decoder, the scalar evaluator, and the shader IR every codegen consumes. Emitters
(SPIR-V, MSL, DXIL) live in sibling SharpEmu.ShaderCompiler.* projects; nothing
here may depend on a host graphics API. -->
<PropertyGroup>
<GenerateDocumentationFile>false</GenerateDocumentationFile>
</PropertyGroup>
<ItemGroup>
<ProjectReference Include="..\SharpEmu.HLE\SharpEmu.HLE.csproj" />
</ItemGroup>
</Project>
@@ -0,0 +1,16 @@
{
"version": 2,
"dependencies": {
"net10.0": {
"sharpemu.hle": {
"type": "Project",
"dependencies": {
"SharpEmu.Logging": "[0.0.1, )"
}
},
"sharpemu.logging": {
"type": "Project"
}
}
}
}