Compare commits

..

8 Commits

Author SHA1 Message Date
Berk ff3ac0bba1 chore: bump version to 0.0.2-beta.3 (#345) 2026-07-17 19:01:35 +03:00
tensorcrush 847371d2de [AGC] Decode VOP3P and emit packed f16 arithmetic (first slice) (#145)
* [AGC] Decode VOP3P and emit packed f16 arithmetic (first slice)

On gfx10 the VOP3P family lives under its own 0b110011000 prefix (word0
top byte 0xCC), which the major-opcode switch currently routes to the
SMEM branch, so packed instructions were decoded as scalar memory ops
and emitted as silent no-ops. Intercept the exact 9-bit prefix ahead of
the switch, decode the five packed-f16 arithmetic opcodes with their
op_sel/op_sel_hi/neg_lo/neg_hi/clamp modifiers, and emit them as
UnpackHalf2x16 -> component-wise f32 vec2 ops -> PackHalf2x16 so no
Float16 capability is needed.

Bit layout and opcode numbers pinned to LLVM MC test encodings
(vop3p.s, gfx10_vop3p_literalv216.txt) and VOP3PInstructions.td.
Unsupported modifiers, packed constants and out-of-scope packed opcodes
fail with a clear error instead of emitting wrong results.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* [AGC] Make packed f16 exact and drop v_pk_fma_f16 (review response)

Address the FP16 correctness review on the VOP3P slice.

Replace GLSL UnpackHalf2x16/PackHalf2x16 with explicit integer f16<->f32
conversions (EmitHalfToFloat/EmitFloatToHalf): exact widening with subnormal
normalisation, and narrowing with round-to-nearest-even, overflow-to-Inf and
NaN/Inf handling. Their subnormal and rounding behaviour no longer depends on
implementation-defined float-controls modes.

With exact conversions, v_pk_add_f16 and v_pk_mul_f16 are bit-exact to a true
f16 op (f32 result rounds losslessly to f16; a f16 product fits in f32). Emit
v_pk_min_f16/v_pk_max_f16 as fminnum_like/fmaxnum_like (NaN operand returns the
other; ordered numeric compare) instead of GLSL FMin/FMax.

v_pk_fma_f16 now fails emission loudly: a fused f16 FMA rounds once, an f32
multiply-add then pack double-rounds (fma(0x4100,0x7522,0x04EA) is 0x7A6B fused
vs 0x7A6A via f32). Exact fused emulation is a planned follow-up slice.

ShaderDump gains an Expect model (Translates/DecodeFails/EmitFails) and packed
regressions: arith, non-default modifiers, and loud-failure pins for the fma
case above and for clamp.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: tensorcrush <tensorcrush@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Antigravity AI <antigravity@gemini.com>
2026-07-17 18:58:22 +03:00
Berk fbafd3f429 [GUI] Add native Vulkan host surface support and more (#337)
* [GUI] Add native Vulkan host surface support and more

* reuse

* [GUI] set width session menu
2026-07-17 18:57:10 +03:00
Mees van den Kieboom aa25f6e978 [Loader] Restore PS5 SELF header support (#342)
Accept both PS4 and PS5 SELF signatures without treating version, key type, or flags as layout markers. Add synthetic coverage for valid header variants and malformed structural fields.

Signed-off-by: missatjuhvdk1 <177474143+missatjuhvdk1@users.noreply.github.com>
Co-authored-by: missatjuhvdk1 <177474143+missatjuhvdk1@users.noreply.github.com>
2026-07-17 18:55:16 +03:00
Raiyan b9eee497aa [VideoOut] Scope the AMD integrated-GPU penalty to Windows (#339)
* [VideoOut] Prefer real integrated GPUs over software rasterizers (#325)

Penalize only AMD integrated GPUs (the #97 vkCreateGraphicsPipelines
crash) instead of all integrated devices, so Intel/Apple/Qualcomm iGPUs
outrank Cpu-type software rasterizers (Mesa lavapipe). Hoist
ScorePhysicalDevice to the outer class and add unit tests for the
ordering.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* [VideoOut] Scope AMD iGPU penalty to Windows via a penalty helper

Extract ComputeDevicePenalty (the value subtracted from a device's base
score) and gate the #97 AMD-integrated penalty on Windows only. Mesa RADV
on Linux (e.g. the Steam Deck's AMD APU) is a different, working driver
and should keep its full integrated score. Add a Steam Deck test case and
drop inline comments.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* [VideoOut] Move device-scoring helpers out of the const block

Relocate ScorePhysicalDevice and ComputeDevicePenalty below the leading
const cluster instead of splitting it, and trim the vendor-ID reference
comment to adapters an x86-64 host can realistically enumerate.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 18:48:09 +03:00
Jose Olguin Lagos a9a4366ef4 test: cover Gen5 decoder fail-closed boundaries (#336)
Reimplement the five synthetic regressions from codefusion-repo/sharpemu#3 and #8 on sharpemu/sharpemu main befd20f415.

Keep the contributor-owned harness and corpus discussed in sharpemu/sharpemu#36 out of scope; use only a focused local recording memory fixture.

Signed-off-by: Jose Olguin Lagos <hellocodefusion@gmail.com>
2026-07-17 18:45:26 +03:00
ParantezTech 76e8f789c4 [github] edit contribution link in pull request template 2026-07-17 18:25:34 +03:00
ParantezTech 6fce0e3055 [github] update PR template and REUSE.toml 2026-07-17 18:20:44 +03:00
11 changed files with 745 additions and 60 deletions
+19 -12
View File
@@ -1,20 +1,27 @@
<!--
Copyright (C) 2026 SharpEmu Emulator Project
SPDX-License-Identifier: GPL-2.0-or-later
-->
## Before submitting
Please read our contribution guidelines before opening a pull request:
➡️ [**CONTRIBUTING.md**](CONTRIBUTING.md)
➡️ [**CONTRIBUTING.md**](https://github.com/sharpemu/sharpemu/blob/main/CONTRIBUTING.md)
By opening this pull request, you confirm that you have read and agree to follow the contribution guidelines.
## Checklist
By submitting this PR, you confirm that:
## Testing
- [ ] I have read `CONTRIBUTING.md`.
- [ ] I tested my changes.
- [ ] I wrote this description myself and did not copy AI-generated explanations.
- [ ] I listed the game(s) I tested, if applicable.
If applicable, list the game(s) you tested and briefly describe the results.
Example:
- Demon's Souls (PPSA01341) Boots to splash screen.
- Dreaming Sarah Save/load works correctly.
If your changes do not affect runtime behavior (e.g. documentation, tooling, CI), write `N/A`.
## Checklist
By submitting this pull request, you confirm that:
- [ ] I have read and followed `CONTRIBUTING.md`.
- [ ] I tested my changes or marked the testing section as `N/A`.
- [ ] I wrote this pull request description myself and did not paste AI-generated explanations.
- [ ] I listed the game(s) I tested (or marked the testing section as `N/A`).
+1 -1
View File
@@ -9,7 +9,7 @@ SPDX-License-Identifier: GPL-2.0-or-later
<ImplicitUsings>enable</ImplicitUsings>
<Nullable>enable</Nullable>
<GenerateDocumentationFile>true</GenerateDocumentationFile>
<SharpEmuVersion>0.0.2-beta.2</SharpEmuVersion>
<SharpEmuVersion>0.0.2-beta.3</SharpEmuVersion>
<Version>$(SharpEmuVersion)</Version>
<RepoRoot>$([MSBuild]::NormalizeDirectory('$(MSBuildThisFileDirectory)'))</RepoRoot>
+1
View File
@@ -10,6 +10,7 @@ path = [
"src/SharpEmu.GUI/Languages/**",
"_logs/**",
".github/images/**",
".github/pull_request_template.md",
"assets/images/**"
]
precedence = "aggregate"
+20 -20
View File
@@ -16,7 +16,9 @@ namespace SharpEmu.Core.Loader;
public sealed class SelfLoader : ISelfLoader
{
private const uint SelfMagic = 0x4F153D1D;
private const uint ElfMagic = 0x7F454C46;
private const uint Ps4SelfMagic = 0x4F153D1D;
private const uint Ps5SelfMagic = 0x5414F5EE;
private const ulong SelfSegmentFlag = 0x800;
private const int PageSize = 0x1000;
private const ulong ImportStubBaseAddress = 0x0000_7000_0000_0000UL;
@@ -355,10 +357,11 @@ public sealed class SelfLoader : ISelfLoader
throw new InvalidDataException("Input image is too small to contain an ELF header.");
}
if (imageData.Length >= sizeof(uint) && BinaryPrimitives.ReadUInt32BigEndian(imageData[..sizeof(uint)]) == SelfMagic)
var magic = BinaryPrimitives.ReadUInt32BigEndian(imageData[..sizeof(uint)]);
if (magic is Ps4SelfMagic or Ps5SelfMagic)
{
var selfHeader = ReadUnmanaged<SelfHeader>(imageData, 0);
if (!selfHeader.HasKnownLayout || selfHeader.Unknown != 0x22)
if (!selfHeader.HasKnownLayout)
{
throw new InvalidDataException("SELF header signature is not recognized.");
}
@@ -381,13 +384,11 @@ public sealed class SelfLoader : ISelfLoader
// acceptable here; anything else — most commonly a still-encrypted
// retail eboot — must be reported clearly rather than failing later
// with an opaque "not a valid ELF header" message.
const uint ElfMagicBigEndian = 0x7F454C46; // "\x7fELF"
var leadingWord = BinaryPrimitives.ReadUInt32BigEndian(imageData[..sizeof(uint)]);
if (leadingWord != ElfMagicBigEndian)
if (magic != ElfMagic)
{
throw new InvalidDataException(
$"Image is neither a decrypted ELF nor a recognized fake-signed SELF " +
$"(leading bytes 0x{leadingWord:X8}). This is almost certainly a still-encrypted " +
$"(leading bytes 0x{magic:X8}). This is almost certainly a still-encrypted " +
$"retail eboot — SharpEmu has no decryption keys and requires a decrypted / " +
$"fake-signed (fSELF) image.");
}
@@ -2789,28 +2790,27 @@ public sealed class SelfLoader : ISelfLoader
private readonly ushort _size2;
private readonly ulong _fileSize;
private readonly ushort _segmentCount;
private readonly ushort _unknown;
private readonly ushort _flags;
private readonly uint _padding;
public ushort SegmentCount => _segmentCount;
public ushort Unknown => _unknown;
public ulong FileSize => _fileSize;
// Version, key type, and flags are signing metadata and vary across
// valid SELF images. They do not change the fixed header layout.
public bool HasKnownLayout =>
_ident0 == 0x4F &&
_ident1 == 0x15 &&
_ident2 == 0x3D &&
_ident3 == 0x1D &&
_ident4 == 0x00 &&
((_ident0 == 0x4F &&
_ident1 == 0x15 &&
_ident2 == 0x3D &&
_ident3 == 0x1D) ||
(_ident0 == 0x54 &&
_ident1 == 0x14 &&
_ident2 == 0xF5 &&
_ident3 == 0xEE)) &&
_ident5 == 0x01 &&
_ident6 == 0x01 &&
_ident7 == 0x12 &&
_ident8 == 0x01 &&
_ident9 == 0x01 &&
_ident10 == 0x00 &&
_ident11 == 0x00;
_ident7 == 0x12;
}
[StructLayout(LayoutKind.Sequential, Pack = 1)]
@@ -295,6 +295,10 @@ internal static unsafe class VulkanVideoPresenter
private static bool _presenterCloseRequested;
private const string DebugUtilsExtensionName = "VK_EXT_debug_utils";
private const uint NvidiaVendorId = 0x10DE;
private const uint AmdVendorId = 0x1002;
// Other GPU PCI vendor IDs, for reference when adding future rules:
// Intel 0x8086, Apple 0x106B, Qualcomm 0x5143 (Windows-on-ARM), Microsoft software 0x1414.
private const int LastResortPenalty = 1000;
private const string PortabilityEnumerationExtensionName = "VK_KHR_portability_enumeration";
private const string PortabilitySubsetExtensionName = "VK_KHR_portability_subset";
private const int GlfwPlatformHint = 0x00050003;
@@ -303,6 +307,50 @@ internal static unsafe class VulkanVideoPresenter
private const int GlfwPlatformWayland = 0x00060003;
private const int GlfwPlatformX11 = 0x00060004;
private const int GlfwPlatformNull = 0x00060005;
internal static int ScorePhysicalDevice(
PhysicalDeviceProperties properties,
string name,
string? deviceOverride)
{
if (!string.IsNullOrWhiteSpace(deviceOverride))
{
return name.Contains(deviceOverride, StringComparison.OrdinalIgnoreCase) ? 1000 : -1000;
}
var score = properties.DeviceType switch
{
PhysicalDeviceType.DiscreteGpu => 300,
PhysicalDeviceType.VirtualGpu => 100,
PhysicalDeviceType.IntegratedGpu => 50,
PhysicalDeviceType.Cpu => 20,
_ => 10,
};
if (properties.VendorID == NvidiaVendorId)
{
score += 500;
}
score -= ComputeDevicePenalty(properties, OperatingSystem.IsWindows());
return score;
}
internal static int ComputeDevicePenalty(PhysicalDeviceProperties properties, bool isWindows)
{
var penalty = 0;
if (isWindows &&
properties.DeviceType == PhysicalDeviceType.IntegratedGpu &&
properties.VendorID == AmdVendorId)
{
penalty += LastResortPenalty;
}
return penalty;
}
private static bool _splashHidden;
private static long _enqueuedGuestWorkSequence;
// Largest contiguous completed sequence, retained for compact diagnostics.
@@ -3518,33 +3566,6 @@ internal static unsafe class VulkanVideoPresenter
}
}
private static int ScorePhysicalDevice(
PhysicalDeviceProperties properties,
string name,
string? deviceOverride)
{
if (!string.IsNullOrWhiteSpace(deviceOverride))
{
return name.Contains(deviceOverride, StringComparison.OrdinalIgnoreCase) ? 1000 : -1000;
}
var score = properties.DeviceType switch
{
PhysicalDeviceType.DiscreteGpu => 300,
PhysicalDeviceType.VirtualGpu => 100,
PhysicalDeviceType.Cpu => 50,
PhysicalDeviceType.IntegratedGpu => -100,
_ => 10,
};
if (properties.VendorID == NvidiaVendorId)
{
score += 500;
}
return score;
}
private void LoadComputeDeviceLimits()
{
_vk.GetPhysicalDeviceProperties(_physicalDevice, out var properties);
@@ -948,6 +948,27 @@ public static partial class Gen5SpirvTranslator
vector);
break;
}
case "VPkAddF16":
case "VPkMulF16":
case "VPkMinF16":
case "VPkMaxF16":
if (!TryEmitPackedF16(instruction, out result, out error))
{
return false;
}
break;
case "VPkFmaF16":
// Deliberately loud: a fused f16 FMA rounds the product+add once,
// whereas doing the multiply-add in f32 and rounding to f16 at the
// end double-rounds. Concrete miss: fma(0x4100, 0x7522, 0x04EA) is
// 0x7A6B fused but 0x7A6A via f32. Exact emulation (round-to-odd
// f32 product then RNE pack) is a planned follow-up slice.
error =
$"unsupported vop3p opcode {instruction.Opcode} " +
"(fused f16 FMA requires single-rounding; deferred to a later slice)";
return false;
default:
error = $"unsupported vector opcode {instruction.Opcode}";
return false;
@@ -976,6 +997,213 @@ public static partial class Gen5SpirvTranslator
return true;
}
// Packed f16 (VOP3P) arithmetic. Each source register holds two f16 values,
// one per result lane. Every f16<->f32 conversion is done with the explicit
// integer sequences below (EmitHalfToFloat / EmitFloatToHalf) instead of
// GLSL UnpackHalf2x16 / PackHalf2x16, whose subnormal and rounding behaviour
// is implementation-defined without float-controls execution modes. The two
// lanes are computed independently: each operand half is widened exactly to
// f32, op_sel/op_sel_hi pick the source half and neg_lo/neg_hi negate it, the
// op runs in f32, and the result is rounded back to f16 with round-to-nearest-
// even. For add and mul this is bit-exact to a true f16 op (the f32 result
// rounds losslessly to f16 by the double-rounding theorem; a f16 product even
// fits in f32 exactly). min/max carry no rounding, so they are exact once the
// conversions are. v_pk_fma_f16 is intentionally not routed here because a
// fused f16 FMA cannot be reproduced by an f32 multiply-add plus a pack.
private bool TryEmitPackedF16(
Gen5ShaderInstruction instruction,
out uint result,
out string error)
{
result = 0;
error = string.Empty;
if (instruction.Control is not Gen5Vop3pControl control)
{
error = $"missing vop3p control for {instruction.Opcode}";
return false;
}
if (control.Clamp)
{
error = $"unsupported vop3p modifiers (clamp) for {instruction.Opcode}";
return false;
}
for (var index = 0; index < 2; index++)
{
var source = instruction.Sources[index];
if (source.Kind is not (Gen5OperandKind.VectorRegister or Gen5OperandKind.ScalarRegister))
{
error =
$"unsupported vop3p operand {source} for {instruction.Opcode} (first slice: registers only)";
return false;
}
}
var low = EmitPackedF16Lane(instruction, control, highLane: false);
var high = EmitPackedF16Lane(instruction, control, highLane: true);
result = BitwiseOr(low, ShiftLeftLogical(high, UInt(16)));
return true;
}
// Computes one result lane (low or high) as a packed 16-bit f16 value.
private uint EmitPackedF16Lane(
Gen5ShaderInstruction instruction,
Gen5Vop3pControl control,
bool highLane)
{
var left = EmitPackedF16Operand(instruction, control, 0, highLane);
var right = EmitPackedF16Operand(instruction, control, 1, highLane);
var value = instruction.Opcode switch
{
"VPkAddF16" => _module.AddInstruction(SpirvOp.FAdd, _floatType, left, right),
"VPkMulF16" => _module.AddInstruction(SpirvOp.FMul, _floatType, left, right),
"VPkMinF16" => EmitPackedF16MinMax(left, right, isMax: false),
"VPkMaxF16" => EmitPackedF16MinMax(left, right, isMax: true),
_ => left,
};
return EmitFloatToHalf(Bitcast(_uintType, value));
}
// Reads source `index`, selects the half feeding this lane (op_sel / op_sel_hi),
// widens it exactly to f32 and applies the lane's negate modifier (neg_lo / neg_hi).
private uint EmitPackedF16Operand(
Gen5ShaderInstruction instruction,
Gen5Vop3pControl control,
int index,
bool highLane)
{
var raw = GetRawSource(instruction, index);
var selectMask = highLane ? control.OpSelHiMask : control.OpSelMask;
var half = ((selectMask >> index) & 1) != 0
? ShiftRightLogical(raw, UInt(16))
: raw;
var value = Bitcast(_floatType, EmitHalfToFloat(half));
var negateMask = highLane ? control.NegHiMask : control.NegLoMask;
if (((negateMask >> index) & 1) != 0)
{
value = _module.AddInstruction(SpirvOp.FNegate, _floatType, value);
}
return value;
}
// fminnum_like / fmaxnum_like: if one operand is NaN return the other; if both
// are NaN return a NaN; otherwise the ordered smaller/larger. The ordering of
// -0/+0 is unspecified under these opcodes, so the ordered compare is enough.
private uint EmitPackedF16MinMax(uint left, uint right, bool isMax)
{
var compare = _module.AddInstruction(
isMax ? SpirvOp.FOrdGreaterThan : SpirvOp.FOrdLessThan,
_boolType,
left,
right);
var numeric = _module.AddInstruction(
SpirvOp.Select, _floatType, compare, left, right);
var leftNan = _module.AddInstruction(SpirvOp.IsNan, _boolType, left);
var rightNan = _module.AddInstruction(SpirvOp.IsNan, _boolType, right);
var withRight = _module.AddInstruction(
SpirvOp.Select, _floatType, rightNan, left, numeric);
return _module.AddInstruction(
SpirvOp.Select, _floatType, leftNan, right, withRight);
}
// Widens an f16 value held in the low 16 bits of `halfBits` to an f32 bit
// pattern, exactly (subnormals normalised, Inf/NaN and signed zero preserved).
// Mirrors the branchless HalfToFloat reference validated against System.Half.
private uint EmitHalfToFloat(uint halfBits)
{
var sign = ShiftLeftLogical(BitwiseAnd(halfBits, UInt(0x8000)), UInt(16));
var exponent = BitwiseAnd(ShiftRightLogical(halfBits, UInt(10)), UInt(0x1F));
var mantissa = BitwiseAnd(halfBits, UInt(0x3FF));
var normal = BitwiseOr(
ShiftLeftLogical(IAdd(exponent, UInt(112)), UInt(23)),
ShiftLeftLogical(mantissa, UInt(13)));
var infinityNan = BitwiseOr(UInt(0x7F80_0000), ShiftLeftLogical(mantissa, UInt(13)));
// Subnormal: normalise the mantissa. FindUMsb of (mantissa | 1) keeps the
// op defined when mantissa is 0; that lane is discarded by the select below.
var highBit = Ext(75, _uintType, BitwiseOr(mantissa, UInt(1)));
var shift = ISubU(UInt(23), highBit);
var subFraction = BitwiseAnd(ShiftLeftLogical(mantissa, shift), UInt(0x7F_FFFF));
var subnormal = SelectU(
IsNotZero(mantissa),
BitwiseOr(ShiftLeftLogical(IAdd(highBit, UInt(103)), UInt(23)), subFraction),
UInt(0));
var magnitude = SelectU(
Equal(exponent, 0),
subnormal,
SelectU(Equal(exponent, 31), infinityNan, normal));
return BitwiseOr(sign, magnitude);
}
// Narrows an f32 bit pattern to an f16 value in the low 16 bits, rounding to
// nearest even (subnormals, overflow-to-Inf and NaN/Inf handled). Mirrors the
// branchless FloatToHalf reference validated exhaustively against System.Half.
private uint EmitFloatToHalf(uint bits)
{
var sign = BitwiseAnd(ShiftRightLogical(bits, UInt(16)), UInt(0x8000));
var absolute = BitwiseAnd(bits, UInt(0x7FFF_FFFF));
var isInfinityNan = UCmp(SpirvOp.UGreaterThanEqual, absolute, UInt(0x7F80_0000));
var isNan = UCmp(SpirvOp.UGreaterThan, absolute, UInt(0x7F80_0000));
var infinityNan = BitwiseOr(
BitwiseOr(sign, UInt(0x7C00)),
SelectU(isNan, UInt(0x200), UInt(0)));
var exponent = ShiftRightLogical(absolute, UInt(23));
var mantissa = BitwiseAnd(absolute, UInt(0x7F_FFFF));
var significand = BitwiseOr(mantissa, UInt(0x80_0000));
// Normal path: round the 24-bit significand down to 11 bits (>> 13) with
// round-to-nearest-even; the carry folds naturally into the exponent.
var roundBit = BitwiseAnd(ShiftRightLogical(significand, UInt(13)), UInt(1));
var rounded = ShiftRightLogical(IAdd(IAdd(significand, UInt(0xFFF)), roundBit), UInt(13));
var halfExponent = ISubU(exponent, UInt(112));
var normalBits = IAdd(ShiftLeftLogical(halfExponent, UInt(10)), ISubU(rounded, UInt(0x400)));
var normal = SelectU(
UCmp(SpirvOp.UGreaterThanEqual, exponent, UInt(113)),
SelectU(UCmp(SpirvOp.UGreaterThanEqual, normalBits, UInt(0x7C00)), UInt(0x7C00), normalBits),
UInt(0));
// Subnormal path: value = round(significand >> (126 - exponent)) with RNE.
// The shift is clamped to 25 so it stays defined; on this path it is >= 14.
var distance = ISubU(UInt(126), exponent);
var shift = SelectU(UCmp(SpirvOp.UGreaterThan, distance, UInt(25)), UInt(25), distance);
var shiftMask = ISubU(ShiftLeftLogical(UInt(1), shift), UInt(1));
var halfWay = ShiftLeftLogical(UInt(1), ISubU(shift, UInt(1)));
var lowBits = BitwiseAnd(significand, shiftMask);
var quotient = ShiftRightLogical(significand, shift);
var roundUp = _module.AddInstruction(
SpirvOp.LogicalOr,
_boolType,
UCmp(SpirvOp.UGreaterThan, lowBits, halfWay),
_module.AddInstruction(
SpirvOp.LogicalAnd,
_boolType,
Equal(lowBits, halfWay),
IsNotZero(BitwiseAnd(quotient, UInt(1)))));
var subnormal = IAdd(quotient, SelectU(roundUp, UInt(1), UInt(0)));
var isSubnormal = UCmp(SpirvOp.ULessThanEqual, exponent, UInt(112));
var finite = SelectU(isSubnormal, subnormal, normal);
return SelectU(isInfinityNan, infinityNan, BitwiseOr(sign, finite));
}
private uint SelectU(uint condition, uint whenTrue, uint whenFalse) =>
_module.AddInstruction(SpirvOp.Select, _uintType, condition, whenTrue, whenFalse);
private uint UCmp(SpirvOp operation, uint left, uint right) =>
_module.AddInstruction(operation, _boolType, left, right);
private uint Equal(uint value, uint constant) =>
_module.AddInstruction(SpirvOp.IEqual, _boolType, value, UInt(constant));
private uint ISubU(uint left, uint right) =>
_module.AddInstruction(SpirvOp.ISub, _uintType, left, right);
private bool TryEmitVectorCompare(
Gen5ShaderInstruction instruction,
out string error)
@@ -237,6 +237,17 @@ public sealed record Gen5SdwaControl(
bool Clamp,
uint? ScalarDestination) : Gen5InstructionControl;
// Packed (VOP3P) source and destination modifiers. Each mask holds one bit per
// source operand. OpSel/OpSelHi pick which 16-bit half of a source feeds the low
// and high result lanes respectively; NegLo/NegHi negate the value routed to each
// lane. Clamp saturates each output half to [0, 1].
public sealed record Gen5Vop3pControl(
uint OpSelMask,
uint OpSelHiMask,
uint NegLoMask,
uint NegHiMask,
bool Clamp) : Gen5InstructionControl;
public sealed record Gen5DppControl(
uint Control,
bool FetchInactive,
@@ -562,6 +562,22 @@ public static class Gen5ShaderTranslator
return DecodeSop(word, out name, out sizeDwords, out error);
}
// gfx10 moved VOP3P (packed 16-bit math) to its own 0b110011000 prefix
// (word0 top byte 0xCC), separate from the VOP3 block. Match the full
// 9-bit prefix here, before the coarse major-opcode switch, so packed
// instructions are not misread as one of the neighbouring encodings.
if ((word & 0xFF800000u) == 0xCC000000u)
{
encoding = Gen5ShaderEncoding.Vop3p;
if (!ctx.TryReadUInt32(baseAddress + pc + sizeof(uint), out var vop3pExtra))
{
error = $"vop3p-extra-read-failed pc=0x{pc:X}";
return false;
}
return DecodeVop3p(word, vop3pExtra, out name, out sizeDwords, out error);
}
switch (word >> 26)
{
case 0x33:
@@ -1160,6 +1176,37 @@ public static class Gen5ShaderTranslator
return true;
}
private static bool DecodeVop3p(
uint word,
uint extra,
out string name,
out uint sizeDwords,
out string error)
{
var opcode = (word >> 16) & 0x7F;
var src0 = extra & 0x1FF;
var src1 = (extra >> 9) & 0x1FF;
var src2 = (extra >> 18) & 0x1FF;
sizeDwords = src0 == 0xFF || src1 == 0xFF || src2 == 0xFF ? 3u : 2u;
error = string.Empty;
// Opcode numbers taken from LLVM's AMDGPU VOP3PInstructions.td and the
// gfx9/gfx10 MC test encodings; they are unchanged across gfx9 and gfx10.
// Unhandled packed opcodes (integer, fma_mix, ...) stay opaque here and
// fail loudly at emission rather than being silently mis-emitted.
name = opcode switch
{
0x0E => "VPkFmaF16",
0x0F => "VPkAddF16",
0x10 => "VPkMulF16",
0x11 => "VPkMinF16",
0x12 => "VPkMaxF16",
_ => $"Vop3pRaw{opcode:X2}",
};
return true;
}
private static bool DecodeDs(
uint word,
out string name,
@@ -1839,6 +1886,28 @@ public static class Gen5ShaderTranslator
isVop3B ? (word >> 8) & 0x7F : null);
break;
}
case Gen5ShaderEncoding.Vop3p:
{
var extra = words[1];
sources =
[
Gen5Operand.Source(extra & 0x1FF, literal),
Gen5Operand.Source((extra >> 9) & 0x1FF, literal),
Gen5Operand.Source((extra >> 18) & 0x1FF, literal),
];
destinations = [Gen5Operand.Vector(word & 0xFF)];
// op_sel_hi is split across both dwords: bits [1:0] live in word1
// [28:27], bit [2] in word0 [14].
var opSelHi = ((extra >> 27) & 0x3) | (((word >> 14) & 0x1) << 2);
control = new Gen5Vop3pControl(
(word >> 11) & 0x7,
opSelHi,
(extra >> 29) & 0x7,
(word >> 8) & 0x7,
((word >> 15) & 1) != 0);
break;
}
case Gen5ShaderEncoding.Ds:
{
var extra = words[1];
@@ -0,0 +1,165 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using System.Buffers.Binary;
using SharpEmu.HLE;
using SharpEmu.ShaderCompiler;
using Xunit;
namespace SharpEmu.Libs.Tests.Agc;
public sealed class Gen5ShaderDecoderBoundaryTests
{
private const ulong ShaderAddress = 0x1_0000_0000;
private const uint Export = 0xF8000000;
private const uint Nop = 0xBF800000;
private const int MaximumInstructionCount = 4096;
[Fact]
public void MissingAddress_IsRejectedWithoutReadingGuestMemory()
{
var memory = new RecordingCpuMemory(ShaderAddress, []);
var decoded = Decode(memory, 0, out var program, out var error);
Assert.False(decoded);
Assert.Empty(program.Instructions);
Assert.Equal("missing", error);
Assert.Empty(memory.Reads);
}
[Fact]
public void UnreadableFirstWord_IsRejectedAtProgramStart()
{
var memory = new RecordingCpuMemory(ShaderAddress, []);
var decoded = Decode(memory, ShaderAddress, out var program, out var error);
Assert.False(decoded);
Assert.Empty(program.Instructions);
Assert.Equal("read-failed pc=0x0", error);
Assert.Collection(
memory.Reads,
read => Assert.Equal(new MemoryRead(ShaderAddress, sizeof(uint), false), read));
}
[Fact]
public void TruncatedTwoDwordInstruction_IsRejectedAtMissingWord()
{
var memory = RecordingCpuMemory.FromWords(ShaderAddress, Export);
var decoded = Decode(memory, ShaderAddress, out var program, out var error);
Assert.False(decoded);
Assert.Empty(program.Instructions);
Assert.Equal("read-failed pc=0x4", error);
Assert.Collection(
memory.Reads,
read => Assert.Equal(new MemoryRead(ShaderAddress, sizeof(uint), true), read),
read => Assert.Equal(
new MemoryRead(ShaderAddress + sizeof(uint), sizeof(uint), false),
read));
}
[Fact]
public void UnknownTopLevelEncoding_IsRejectedWithoutSpeculativeReads()
{
const uint unknownTopLevel = 0xE4000000;
var memory = RecordingCpuMemory.FromWords(ShaderAddress, unknownTopLevel);
var decoded = Decode(memory, ShaderAddress, out var program, out var error);
Assert.False(decoded);
Assert.Empty(program.Instructions);
Assert.Equal("unknown-top pc=0x0 word=0xE4000000", error);
Assert.Collection(
memory.Reads,
read => Assert.Equal(new MemoryRead(ShaderAddress, sizeof(uint), true), read));
}
[Fact]
public void MaximumInstructionCountWithoutEndPgm_IsRejectedAsUnterminated()
{
var words = new uint[MaximumInstructionCount];
Array.Fill(words, Nop);
var memory = RecordingCpuMemory.FromWords(ShaderAddress, words);
var decoded = Decode(memory, ShaderAddress, out var program, out var error);
Assert.False(decoded);
Assert.Empty(program.Instructions);
Assert.Equal("unterminated", error);
Assert.Equal(MaximumInstructionCount, memory.Reads.Count);
Assert.All(memory.Reads, read => Assert.True(read.Succeeded));
Assert.Equal(
new MemoryRead(
ShaderAddress + ((MaximumInstructionCount - 1) * sizeof(uint)),
sizeof(uint),
true),
memory.Reads[^1]);
}
private static bool Decode(
RecordingCpuMemory memory,
ulong address,
out Gen5ShaderProgram program,
out string error) =>
Gen5ShaderTranslator.TryDecodeProgram(
new CpuContext(memory, Generation.Gen5),
address,
out program,
out error);
private readonly record struct MemoryRead(ulong Address, int Length, bool Succeeded);
private sealed class RecordingCpuMemory(ulong baseAddress, byte[] storage) : ICpuMemory
{
public List<MemoryRead> Reads { get; } = [];
public static RecordingCpuMemory FromWords(ulong baseAddress, params uint[] words)
{
var storage = new byte[words.Length * sizeof(uint)];
for (var index = 0; index < words.Length; index++)
{
BinaryPrimitives.WriteUInt32LittleEndian(
storage.AsSpan(index * sizeof(uint), sizeof(uint)),
words[index]);
}
return new RecordingCpuMemory(baseAddress, storage);
}
public bool TryRead(ulong virtualAddress, Span<byte> destination)
{
var succeeded = TryResolve(virtualAddress, destination.Length, out var offset);
Reads.Add(new MemoryRead(virtualAddress, destination.Length, succeeded));
if (succeeded)
{
storage.AsSpan(offset, destination.Length).CopyTo(destination);
}
return succeeded;
}
public bool TryWrite(ulong virtualAddress, ReadOnlySpan<byte> source) => false;
private bool TryResolve(ulong virtualAddress, int length, out int offset)
{
offset = 0;
if (virtualAddress < baseAddress)
{
return false;
}
var relative = virtualAddress - baseAddress;
if (relative > (ulong)storage.Length ||
(ulong)length > (ulong)storage.Length - relative)
{
return false;
}
offset = (int)relative;
return true;
}
}
}
@@ -0,0 +1,89 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using System.Buffers.Binary;
using SharpEmu.Core.Loader;
using SharpEmu.Core.Memory;
using Xunit;
namespace SharpEmu.Libs.Tests.Loader;
public sealed class SelfLoaderTests
{
private const uint Ps4SelfMagic = 0x4F153D1D;
private const uint Ps5SelfMagic = 0x5414F5EE;
private const int SelfHeaderSize = 0x20;
private const int ElfHeaderSize = 0x40;
[Theory]
[InlineData(Ps4SelfMagic, (byte)0x00, 0x0000_0101u, (ushort)0x22)]
[InlineData(Ps5SelfMagic, (byte)0x00, 0x0000_0101u, (ushort)0x22)]
[InlineData(Ps5SelfMagic, (byte)0x10, 0x1000_0101u, (ushort)0x32)]
public void Load_AcceptsSupportedSelfHeaderVariants(
uint magic,
byte version,
uint keyType,
ushort flags)
{
var imageData = CreateSelfImage(magic, version, keyType, flags);
var image = new SelfLoader().Load(imageData, new VirtualMemory());
Assert.True(image.IsSelf);
Assert.Equal(2, image.ElfHeader.AbiVersion);
Assert.Empty(image.ProgramHeaders);
Assert.Empty(image.MappedRegions);
}
[Theory]
[InlineData(0x05, (byte)0x00)]
[InlineData(0x06, (byte)0x02)]
[InlineData(0x07, (byte)0x00)]
public void Load_RejectsUnsupportedStructuralSelfHeaderValues(int offset, byte value)
{
var imageData = CreateSelfImage(Ps5SelfMagic, 0x10, 0x1000_0101, 0x32);
imageData[offset] = value;
Assert.Throws<InvalidDataException>(() =>
new SelfLoader().Load(imageData, new VirtualMemory()));
}
private static byte[] CreateSelfImage(uint magic, byte version, uint keyType, ushort flags)
{
var imageData = new byte[SelfHeaderSize + ElfHeaderSize];
var selfHeader = imageData.AsSpan(0, SelfHeaderSize);
BinaryPrimitives.WriteUInt32BigEndian(selfHeader, magic);
selfHeader[0x04] = version;
selfHeader[0x05] = 0x01;
selfHeader[0x06] = 0x01;
selfHeader[0x07] = 0x12;
BinaryPrimitives.WriteUInt32LittleEndian(selfHeader[0x08..], keyType);
BinaryPrimitives.WriteUInt16LittleEndian(selfHeader[0x0C..], SelfHeaderSize);
BinaryPrimitives.WriteUInt64LittleEndian(selfHeader[0x10..], (ulong)imageData.Length);
BinaryPrimitives.WriteUInt16LittleEndian(selfHeader[0x18..], 0);
BinaryPrimitives.WriteUInt16LittleEndian(selfHeader[0x1A..], flags);
WriteMinimalElfHeader(imageData.AsSpan(SelfHeaderSize, ElfHeaderSize));
return imageData;
}
private static void WriteMinimalElfHeader(Span<byte> header)
{
header.Clear();
header[0x00] = 0x7F;
header[0x01] = (byte)'E';
header[0x02] = (byte)'L';
header[0x03] = (byte)'F';
header[0x04] = 2;
header[0x05] = 1;
header[0x06] = 1;
header[0x07] = 9;
header[0x08] = 2;
BinaryPrimitives.WriteUInt16LittleEndian(header[0x10..], 3);
BinaryPrimitives.WriteUInt16LittleEndian(header[0x12..], 62);
BinaryPrimitives.WriteUInt32LittleEndian(header[0x14..], 1);
BinaryPrimitives.WriteUInt64LittleEndian(header[0x20..], ElfHeaderSize);
BinaryPrimitives.WriteUInt16LittleEndian(header[0x34..], ElfHeaderSize);
BinaryPrimitives.WriteUInt16LittleEndian(header[0x36..], 0x38);
}
}
@@ -0,0 +1,94 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using SharpEmu.Libs.VideoOut;
using Silk.NET.Vulkan;
using Xunit;
namespace SharpEmu.Libs.Tests.VideoOut;
public sealed class VulkanPhysicalDeviceScoringTests
{
private const uint NvidiaVendorId = 0x10DE;
private const uint AmdVendorId = 0x1002;
private const uint IntelVendorId = 0x8086;
private const uint QualcommVendorId = 0x5143;
private static PhysicalDeviceProperties Device(PhysicalDeviceType type, uint vendorId = 0)
=> new() { DeviceType = type, VendorID = vendorId };
private static int Score(PhysicalDeviceType type, uint vendorId = 0)
=> VulkanVideoPresenter.ScorePhysicalDevice(Device(type, vendorId), name: string.Empty, deviceOverride: null);
[Fact]
public void RealIntegratedGpuOutranksSoftwareRasterizer()
{
Assert.True(Score(PhysicalDeviceType.IntegratedGpu) > Score(PhysicalDeviceType.Cpu));
}
[Theory]
[InlineData(IntelVendorId)]
[InlineData(QualcommVendorId)]
[InlineData(0u)]
public void NonAmdIntegratedGpuBeatsSoftware(uint vendorId)
{
Assert.True(Score(PhysicalDeviceType.IntegratedGpu, vendorId) > Score(PhysicalDeviceType.Cpu));
}
[Fact]
public void DiscreteGpuOutranksIntegratedGpu()
{
Assert.True(Score(PhysicalDeviceType.DiscreteGpu) > Score(PhysicalDeviceType.IntegratedGpu));
}
[Fact]
public void NvidiaDiscreteGpuScoresHighest()
{
var nvidia = Score(PhysicalDeviceType.DiscreteGpu, NvidiaVendorId);
Assert.True(nvidia > Score(PhysicalDeviceType.DiscreteGpu));
Assert.True(nvidia > Score(PhysicalDeviceType.IntegratedGpu));
}
[Fact]
public void DeviceOverridePinsMatchingAdapterAndExcludesOthers()
{
var properties = Device(PhysicalDeviceType.Cpu);
Assert.Equal(1000, VulkanVideoPresenter.ScorePhysicalDevice(properties, "Radeon Graphics", "radeon"));
Assert.Equal(-1000, VulkanVideoPresenter.ScorePhysicalDevice(properties, "Radeon Graphics", "nvidia"));
}
[Fact]
public void AmdIntegratedGpuIsLastResortOnWindows()
{
var penalty = VulkanVideoPresenter.ComputeDevicePenalty(
Device(PhysicalDeviceType.IntegratedGpu, AmdVendorId), isWindows: true);
Assert.True(penalty > 0);
Assert.True(50 - penalty < 10);
}
[Fact]
public void AmdApuIsNotPenalizedOffWindows()
{
var penalty = VulkanVideoPresenter.ComputeDevicePenalty(
Device(PhysicalDeviceType.IntegratedGpu, AmdVendorId), isWindows: false);
Assert.Equal(0, penalty);
}
[Theory]
[InlineData(true)]
[InlineData(false)]
public void AmdDiscreteGpuIsNeverPenalized(bool isWindows)
{
Assert.Equal(0, VulkanVideoPresenter.ComputeDevicePenalty(
Device(PhysicalDeviceType.DiscreteGpu, AmdVendorId), isWindows));
}
[Theory]
[InlineData(IntelVendorId)]
[InlineData(QualcommVendorId)]
public void NonAmdIntegratedGpuIsNeverPenalized(uint vendorId)
{
Assert.Equal(0, VulkanVideoPresenter.ComputeDevicePenalty(
Device(PhysicalDeviceType.IntegratedGpu, vendorId), isWindows: true));
}
}