* [AGC] Decode gfx10 SOPP hint instructions
s_clause (0x21), s_waitcnt_depctr (0x23), s_round_mode (0x24) and s_denorm_mode (0x25) were missing from the SOPP decode table, so any shader containing one of these scheduling/mode hints failed to decode entirely with unknown-sopp. No emitter changes are needed: non-branch SOPP instructions are already emitted as no-ops. Opcodes verified against LLVM SOPInstructions.td (SOPP_Real_32_gfx10); decode and end-to-end SPIR-V compilation verified with a synthetic program containing all four hints.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* [AGC] Narrow SOPP additions to scheduling hints only
Per review: s_round_mode (0x24) and s_denorm_mode (0x25) write the
shader floating-point MODE state, and the emitter's blanket SOPP no-op
would have silently ignored their simm16 payloads, trading a loud
decode failure for a potential floating-point semantics mismatch. They
are removed and keep failing decode explicitly until their semantics
are modeled or conservatively validated.
s_clause (0x21) and s_waitcnt_depctr (0x23) remain: they are pure
scheduler/dependency hints with no value semantics.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Two gaps around the VOP3 signed multiplies caused whole-shader SPIR-V
compilation failures:
- v_mul_lo_i32 (0x16B) decoded correctly but had no emission case, so
any shader containing it failed with "unsupported vector opcode
VMulLoI32". Its low 32 result bits are identical to the unsigned
multiply in two's complement, so it now shares the v_mul_lo_u32 IMul
case.
- v_mul_hi_i32 (0x16C) was missing from the VOP3 decode table entirely
and decoded as an opaque Vop3Raw16C, which also fails at emission.
It is now decoded and emitted by sign-extending both operands to
64 bits, multiplying, and taking the upper 32 bits of the product,
mirroring the existing v_mul_hi_u32 pattern.
Opcode numbers verified against LLVM's AMDGPU backend
(VOP3Instructions.td): V_MUL_LO_U32 gfx10 = 0x169, V_MUL_HI_U32 =
0x16a, V_MUL_LO_I32 = 0x16b, V_MUL_HI_I32 = 0x16c. Behavior verified
by decoding and fully compiling a synthetic program containing all
four multiplies: previously the 0x16C word decoded as Vop3Raw16C and
compilation failed at the v_mul_lo_i32 instruction; now all four
decode by name and the program compiles to SPIR-V.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
VOP2 opcode 0x2B was mapped to v_ldexp_f32, which is its gfx6/gfx7
assignment. On gfx10-class hardware 0x2B is v_fmac_f32, so any shader
using it silently computed ldexp(a, b) instead of dst += a * b.
v_ldexp_f32 on gfx10 only exists as VOP3 0x362, which the VOP3 table
already maps correctly.
Also add the remaining members of the fmac family:
- v_fmamk_f32 (0x2C) and v_fmaak_f32 (0x2D), including their mandatory
literal dword in instruction sizing and operand construction, reusing
the existing v_madmk/v_madak handling.
- The VOP3-encoded form of v_fmac_f32 (0x12B), emitted when source
modifiers are present.
SPIR-V emission reuses the existing v_mac_f32 body (fma with the
destination register as addend) and the v_mad/v_fma case group.
Opcode assignments verified against LLVM's AMDGPU backend
(VOP2Instructions.td): V_FMAC_F32 gfx10 = 0x02b, V_FMAMK_F32 = 0x02c,
V_FMAAK_F32 = 0x02d; V_LDEXP_F32 is 0x02b only on gfx6/gfx7 and is
VOP3-only 0x362 on gfx10. Decode verified by feeding hand-assembled
gfx1013 words through Gen5ShaderTranslator: 0x560A0501 previously
decoded as VLdexpF32 and a v_fmamk_f32 program failed with
unknown-vop2 op=0x2C; both now decode correctly, and VOP3 0x362 still
decodes as VLdexpF32.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Treat the exact untextured transparent-black premultiplied fill used by Chowdren as an overwrite. This prevents Dreaming Sarah fog and vignette render targets from accumulating across frames; SHARPEMU_DISABLE_TRANSPARENT_FILL_CLEAR=1 restores prior behavior.
* [agc] Add shader type 4 (GS) and register defaults v13 support
Astro Bot (#11) crashes on boot due to two missing GPU features:
1. Shader type 4 (Geometry Shader) — SPI_SHADER_PGM_LO/HI register
offsets 0x8A/0x8B were missing. Added constants and switch cases
for shader type 4 in GetExpectedSpiShaderPgmLo/Hi. Also added
type 4 to IsEsGeometryShaderType (2 or 4 or 6).
2. Register defaults version 13 — was not recognized as supported.
Added RegisterDefaultsVersion13 constant and included it in
IsSupportedRegisterDefaultsVersion.
* [kernel] Add POSIX pthread_cond_timedwait export
SILENT HILL (#4) and Poppy Playtime (#3) crash on boot due to
missing POSIX pthread_cond_timedwait (NID 27bAgiJmOh0).
The Sony wrapper scePthreadCondTimedwait (NID BmMjYxmew1w) was
already implemented, but the raw POSIX symbol was not exported.
Added [SysAbiExport] for pthread_cond_timedwait delegating to
existing PthreadCondWaitCore with timed: true.
* [memory] Fix FlushInstructionCache null process handle
PhysicalVirtualMemory.cs called FlushInstructionCache with null as the
process handle in two places (SetProtection and TryWriteExclusive).
On Windows, a null handle does not reliably resolve to the current
process — the correct call is GetCurrentProcess() (pseudo-handle -1).
Also corrected the P/Invoke signature:
- Changed return type from void to bool with [return: MarshalAs(Bool)]
- Added SetLastError = true
- Added GetCurrentProcess() P/Invoke import
This matches the pattern already used in DirectExecutionBackend.cs
which correctly passes GetCurrentProcess() to all FlushInstructionCache
calls.
* [hle] Distinguish NOT_FOUND from NOT_IMPLEMENTED and log duplicate NIDs
Three diagnostic improvements to the HLE dispatch path:
1. ModuleManager.RegisterFromAssembly — duplicate NID registration was
silently skipped (dispatchTable first-wins, exportTable last-wins,
causing metadata divergence). Now logs a warning with the NID and
export name so conflicts are visible.
2. ModuleManager.TryDispatch — generation mismatch returned
ORBIS_GEN2_ERROR_NOT_FOUND, conflating 'function does not exist'
with 'function exists but not for this generation'. Now returns
ORBIS_GEN2_ERROR_NOT_IMPLEMENTED for generation mismatch, matching
the existing convention in CpuDispatcher. Also adds debug logging
for both NOT_FOUND and NOT_IMPLEMENTED paths.
3. DirectExecutionBackend.Imports.cs — the import dispatch else-branch
(the actual hot path that bypasses ModuleManager.TyDispatch via
cached export) had the same conflation. Split into:
- else if (export exists but generation mismatch) → NOT_IMPLEMENTED
- else (no export at all) → NOT_FOUND
This makes runtime diagnostics correctly distinguish missing exports
from generation-unsupported exports.
* [cpu] Check VirtualProtect return values in all stub creation paths
9 VirtualProtect calls in DirectExecutionBackend.cs had unchecked
return values. If VirtualProtect silently fails, memory protection
remains incorrect — stubs allocated with PAGE_EXECUTE_READWRITE (0x40)
never get downgraded to PAGE_EXECUTE_READ (0x20), or guest thread
entry stubs never get upgraded to writable. This causes access
violations on next execution or silent data corruption.
Fixed all 9 sites with proper error handling:
- 6 stub creation methods (return 0 on failure + log error)
- 2 guest thread entry methods (set reason + return Exception)
- 1 guest entry method (set LastError + return MEMORY_FAULT)
Stub creation sites fixed:
- CreateImportDispatchStub (line ~1683)
- EnsureTlsHandler (void, log + return)
- CreateUnresolvedReturnStub (return 0)
- CreateGuestReturnStub (return 0)
- CreateExceptionHandlerTrampoline (return 0)
- CreateTlsStoreHelperStub (return 0)
Guest thread entry sites fixed:
- StartGuestThreadNativeCall (return Exception)
- StartGuestContinuationNativeCall (return Exception)
- RunGuestEntryPoint (return MEMORY_FAULT)
* [kernel] Remove unused duplicate _nextFileDescriptor field
KernelExports.cs declared _nextFileDescriptor but never used it.
The actual field used for file descriptor allocation lives in
KernelMemoryCompatExports.cs (lines 1314, 1337). This was a dead
duplicate causing CS0414 warning.
Build is now 0 errors, 0 warnings.
---------
Co-authored-by: Hermes Atlas <hermesatlas@example.com>
* [shader-decoder-part1] Implemented a shader decoder for Gen5 shaders, including IR generation, metadata reading, scalar evaluation, and SPIR-V translation. Updated related exports and video output components to support the new shader decoding functionality.
* [shader decoder] correct RDNA2 operands, fixing synchronization problems
* [shader-decoder] RDNA2 decoder improvements
* [shader-decoder] fix RDNA2 shift masking and sprite draws
* [shader-decoder] improve RDNA2 shader decoder to support more instructions and fix some issues with the previous implementation.