mirror of
https://github.com/par274/sharpemu.git
synced 2026-07-22 19:06:15 +08:00
98c6851840a1342c1aa8c5a12ac1e10fd23b6554
58 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0f224ec036 | Gui Settings Null list Entries (#430) | ||
|
|
85dc98dedc | test: add Fiber exports contract tests (13 tests) (#428) | ||
|
|
a60bfc9c83 | [Kernel] Implement pthread semaphore exports (#424) | ||
|
|
09812600a0 | Add libc heap trace contract tests (#409) | ||
|
|
a030cb5a5d |
Gpu runtime stalls (#410)
* [runtime] restore default GC mode * [cpu] add string leaf stubs * [ampr] allow concurrent reads * [bink] keep guest decode path * [kernel] streamline host memory access * [shader] add scalar memory fallback * [gpu] bound guest data pool * [gpu] reduce queue stalls * [video] stabilize guest resources * revert lock file |
||
|
|
336286e588 |
CPU: scan final TLS access pattern offset (#414)
Co-authored-by: Dafenx <196083014+Dafenxz0@users.noreply.github.com> |
||
|
|
bab965e394 | [HLE] Add RandomExports HLE (#413) | ||
|
|
daaeb6213e |
Fix Massive Bug Preventing UE5 Titles From Booting (#406)
* Fix cross platform memcpy bug |
||
|
|
94153955b0 |
[Gpu] Metal backend: complete IGuestGpuBackend implementation on AppKit + Metal (#283)
* [ShaderCompiler.Metal] MSL translator core: dispatcher, EXEC model, compute stage
The Metal codegen backend, rebuilt on the merged backend-neutral
abstractions (replacing the pre-abstraction spike): consumes
(Gen5ShaderState, Gen5ShaderEvaluation) and emits MSL text; the renderer
owns MTLLibrary compilation, mirroring the emitters-produce-bytes rule
the Vulkan sibling documents.
The execution model mirrors Gen5SpirvTranslator: one invocation per GCN
lane (wave32 — natively the Apple simdgroup width), a typeless uint
register file with as_type<float> bitcasts, EXEC/VCC as per-lane bools
whose guest-visible mask registers materialize via simd_ballot, and the
same PC-dispatcher loop over basic blocks with the
SHARPEMU_SHADER_MAX_STEPS iteration guard and the dominating-scalar-
definition dataflow for buffer binding resolution. Unlike SPIR-V, MSL
permits shared prelude functions, so unaligned/subdword buffer access
is a range-checked device-uchar* helper instead of per-site inlining;
buffer byte lengths and the compute dispatch limit travel in one
reserved SharpEmuUniforms constant buffer (Metal has no OpArrayLength).
This first slice covers the compute entry point end to end: scalar/
vector ALU core (moves, int/float arithmetic, FMA family, shifts,
bitfield ops, min/max/med3, conversions, transcendentals with the Tau
scale on sin/cos), the full VCmp/VCmpx compare matrix writing VCC/EXEC,
the saveexec family, scalar compares and SCC-updating SOP2 forms, lane
ops (readfirstlane, mbcnt), VOP3 abs/neg/clamp/omod modifiers, scalar
memory, and raw global/buffer loads, stores, and atomics with EXEC
guards. Unsupported opcodes fail loudly with pc + mnemonic. SDWA/DPP,
typed format loads, LDS, images, and the pixel/vertex stages follow in
the next phases.
Tests live in their own self-contained project (the per-backend model:
depends only on the codegen under test): hand-assembled synthetic
fixtures drive the real decoder end to end, structural assertions and
golden-MSL comparisons run on every platform since translation is pure
text generation, and goldens regenerate via SHARPEMU_UPDATE_GOLDENS=1.
* [ShaderCompiler.Metal] Real-device runtime tests: compile + execute on the GPU
Lifts the spike's objc_msgSend LibraryImport harness (MTLDevice /
MTLCompileOptions with fast-math off, as a real Metal backend must
compile) onto the new translator contract: the guest data buffer binds
at index 0 and the SharpEmuUniforms constant buffer (dispatch limit +
buffer byte lengths) at index 1.
Three runtime tiers on hosts with a Metal device (no-op elsewhere so
Windows/Linux CI stays green): every fixture's emitted MSL must be
accepted by the OS runtime Metal compiler; the exec-store program must
produce bit-exact GPU results including the EXEC-masked store that must
not land; and a scalar countdown loop must iterate through the PC
dispatcher (s_cmp_lg_u32 + s_cbranch_scc1 across five round trips).
The loop fixture also fixes its own hand-assembly: s_sub_i32 sets SCC
to signed overflow, not result-nonzero, so the loop condition uses an
explicit compare.
* [ShaderCompiler.Metal] Phase 2: scalar/vector ALU parity with the SPIR-V translator
Ports the remaining ALU semantics from Gen5SpirvTranslator.Alu so the two
codegens cannot disagree on instruction behavior:
- Carry/borrow family (v_add_co/_ci, v_sub_co/_rev, v_subb/_rev) with the
carry mask written to the VOP3 scalar destination or VCC, ANDed with
EXEC; v_mad_u64_u32 with the 64-bit pair result and carry-out.
- Full SDWA support: byte/word source selects with sign-extension,
integer abs/neg modifiers, and destination-select merge (zero-fill,
sign-extend, preserve) into the previous register value.
- DPP16/DPP8: quad permute, row shl/shr/ror, mirror/half-mirror,
broadcast and xor controls via simd_shuffle, bound-control and
row/bank write-enable masks, fetch-inactive handling; DPP-predicated
compares merge into VCC.
- Lane ops: readfirstlane from the first EXEC-active lane via
ballot+ctz, readlane/writelane, permlane16/permlanex16.
- 64-bit scalar ops over SGPR pairs in real ulong arithmetic (logic
family, shifts, bfe/bfm with width clamping, wqm quad expansion,
cselect, mov, s_getpc) plus the B32/B64 saveexec families.
- Sopk forms decode the signed 16-bit immediate and s_cmpk compares the
destination register; SOPC scalar compares including s_bitcmp0/1.
- Conversions: f16<->f32 via as_type<half>, pkrtz with round-to-zero
mantissa truncation, pknorm via pack_float_to_{s,u}norm2x16, pk_u8
byte insert, off_f32_i4 table, rpi/flr rounding; cube id/sc/tc/ma
decision trees; v_cmp_class_f32; VCCZ/EXECZ/SCC readable as data.
Also fixes a real phase-1 bug the reference surfaced: fmamk/fmaak
sources arrive in natural order from the decoder, so all MAD/FMA forms
are fma(src0, src1, src2) — the previous operand swap computed
v1*v2+K for v_fmamk (should be v1*K+v2). The regenerated fmac golden
shows the corrected expansion, and mirroring the SPIR-V translator,
v_mul_u32_u24 is a full 32-bit multiply (only the hi/mad forms mask).
All 13 Metal tests pass including the real-GPU execution tier.
* [ShaderCompiler.Metal] Phase 3: typed format loads, LDS, and D16 subdword memory
Typed MUBUF/MTBUF loads convert through the descriptor's GFX10 unified
format at execution time, mirroring the SPIR-V translator: the prelude
bakes a 128-entry format table from the shared Gfx10UnifiedFormat
decoder (compiled shaders may be reused with new SRDs, so decoding must
stay dynamic), per-component layouts for the legacy DATA_FORMAT values
drive range-checked unaligned loads, NUM_FORMAT conversion handles
unorm/snorm (clamped at -1)/uscaled/sscaled/uint/sint/float including
f16 and the 10/11-bit unsigned mini-floats of 10_11_11 / 11_11_10, the
missing-component default is one in the format's domain, and dst_sel
swizzling comes from descriptor word 3. Format stores stay raw dword
stores like the reference.
LDS lands as 32 KB of threadgroup memory (gated on the program actually
using DS ops so occupancy is not taxed): ds_read/write b32/b64/b96/b128,
the write2/read2 pairs including st64 scaling, and ds_add_u32 as a
relaxed threadgroup atomic, with EXEC-guarded writes and the address
masked into bounds. Subdword loads/stores gain the D16/D16Hi variants
that merge into one half of the destination register (and shift the
source for high stores), classified the same way as the reference.
New GPU-executed fixture: an LDS round trip (write literal, s_barrier,
read back, store to the buffer) passes bit-exact on a real Metal device
alongside the existing tiers.
* [ShaderCompiler.Metal] Phase 4: pixel stage, images, and interpolation
The pixel entry points land with the same contract as the SPIR-V
translator (single-target and MRT forms, validated for unique guest
slots and dense host locations): the emitted fragment function takes a
stage_in struct carrying [[position]] plus the interpolated attributes
discovered from the program's V_INTERP controls, writes an output
struct with one [[color(hostLocation)]] attachment per binding typed by
its Float/Sint/Uint kind, seeds pixel-input VGPRs in SPI_PS_INPUT_ADDR
compact order from the fragment coordinate, keeps EXEC masking through
translation, and discards lanes that exit with EXEC off. Exports write
MRT targets per component under EXEC (disabled components keep their
previous value) including compressed half-pair exports; vertex-target
exports no-op until the vertex stage.
Images arrive as texture2d<float|int|uint> arguments (storage bindings
as access::read_write) with samplers alongside, classified from the
descriptor's unified format via the shared Gfx10UnifiedFormat decoder
and resolved per instruction with the same dominating-scalar-definition
scheme as buffers. The sample matrix covers implicit LOD, SampleL/Lz,
SampleB, SampleD gradients, PCF compare (manual reference<=texel,
broadcast r,r,r,1), per-lane texel offsets folded into normalized
coordinates by the selected mip extent (Metal sample offsets must be
constants), gather4 including compare and offset forms, clamped
ImageLoad/Mip, bounds-checked EXEC-guarded ImageStore, GetResinfo, and
A16 packed addresses / D16 packed data in both directions.
Graphics stages model LDS as per-invocation scratch (the SPIR-V
Private-array trick) instead of threadgroup memory. Wave ops keep the
invocation's real simdgroup in every stage — Apple fragment simdgroups
make that the same model as compute, where the SPIR-V translator
instead emulates a single logical lane; both round-trip EXEC masks
consistently.
The pixel fixture (interpolated attr0.xy plus inline constants exported
to MRT0) is golden-pinned, structurally asserted, and accepted by the
OS Metal compiler on a real device.
* [ShaderCompiler.Metal] Phase 5: vertex stage and fixed presenter shaders
The vertex entry point completes the four-entry-point contract: the
emitted vertex function takes fetched attributes as a stage_in struct
([[attribute(location)]], bound by the backend via MTLVertexDescriptor
from the reflected vertex inputs), returns [[position]] plus one
[[user(locnN)]] param output per export target 32..63 — unioned with
requiredVertexOutputCount so Metal's exact vertex-out/fragment-in
interface match succeeds, with unexported locations zero-filled —
seeds v5/v8 from [[vertex_id]]/[[instance_id]], intercepts buffer loads
the evaluator captured as fixed-function vertex inputs, and applies the
same EXEC-selected component rules to position/param exports (disabled
components default to 0,0,0,1) including compressed half pairs.
MSL vertex functions have no simdgroup attributes, so the vertex stage
models a single logical wave lane exactly like the SPIR-V translator's
graphics path: lane 0, ballot degrades to 0/1, and lane-shuffle ops
would fail Metal compilation loudly (no real guest vertex shader uses
them).
MslFixedShaders mirrors SpirvFixedShaders for the presenter surface:
the fullscreen-triangle vertex stage (position from the vertex index,
screen-space UV broadcast to every requested attribute location), the
copy/solid/attribute diagnostic fragments, and the output-free
depth-only fragment. Metal forbids "main", so each carries a stable
entry name.
The vertex fixture (constant position + one param export) is
golden-pinned and structurally asserted; it and all five fixed shaders
are accepted by the OS Metal compiler on a real device.
* [ShaderCompiler.Metal] Author static MSL blocks as template files
The prelude helpers (buffer access, ballot, tables), the format-load
conversion functions, and all five fixed presenter shaders move out of
AppendLine walls into Templates/*.msl embedded resources — real Metal
source with syntax highlighting and reviewable diffs — rendered by a
small {{placeholder}} substituter that fails loudly on any
unsubstituted token. Substitution points are deliberately few: the
stage-dependent ballot expression (vertex has no simdgroup attributes),
the baked GFX10 format table and layout cases, and the fixed shaders'
parameters. Per-instruction body emission stays programmatic, where a
template cannot express it.
Behavior-identical by construction: the golden files are untouched and
the whole suite — including the real-device execution and compile
tiers over the templated output — passes against them unchanged. The
.msl files carry no license headers (they would leak into every emitted
shader), so REUSE.toml annotates the Templates directory instead.
* [ShaderCompiler.Metal] Cover the MSL goldens in REUSE.toml
The golden files are verbatim emitter output regenerated by the test
suite; license headers inside them would either break the byte-exact
comparison or force the emitter to write SPDX text into every shader.
Annotate the Goldens directory like the Templates one.
* [ShaderCompiler.Metal] Address review: dominating-binding parity and harness binding indices
The buffer-binding fallback now mirrors the SPIR-V translator: a candidate
binding is accepted only when the descriptor registers hold the exact same
scalar definitions at the target PC as at one of the binding's own access
points (HasSameScalarDefinitions), instead of merely being non-conflicting
at the target. Resolutions are cached per PC like the reference.
The runtime test harness no longer hardcodes buffer indices 0/1 and a
20-byte uniforms blob: TryExecuteSingleThread takes the data/uniforms bind
indices, and ExecuteOrThrow derives them plus the uniforms size from the
compiled shader's GlobalMemoryBindings per the translator contract.
* [Gpu] Add the Metal guest-GPU backend: shader compilation and formats
First phase of the Metal backend behind the IGuestGpuBackend seam. The
backend compiles all three shader stages through Gen5MslTranslator and
exposes the guest render-target format table (mirroring the Vulkan table
case for case; guest format 9 maps to BGR10A2, the Metal layout matching
Vulkan's A2R10G10B10 pack). Wave64 compute is rejected with a clear error
until the two-pass emulation exists.
SHARPEMU_GPU_BACKEND=metal opts in on macOS; Vulkan stays the default on
every platform until the Metal presenter reaches parity. Presenter-side
methods fail loudly instead of dropping guest frames silently.
* [Gpu] Add the Metal presenter core: AppKit window, CAMetalLayer, CPU-frame path
The presenter opens an NSWindow hosting a CAMetalLayer and drives a manually
pumped NSApplication event loop, structured like the Vulkan presenter's
poll-and-render loop and posted onto HostMainThread the same way (AppKit
traps off the process main thread). All OS access goes through objc_msgSend
LibraryImport bindings declared locally — no windowing or binding packages on
this path, which is what keeps it NativeAOT-clean. Struct-returning ObjC
calls are avoided entirely so one calling convention works under Rosetta.
Presents CPU-produced BGRA frames and the splash through a fullscreen
triangle with a dedicated present fragment stage that flips V: with Metal's
y-up NDC the shared fullscreen triangle puts UV (0,0) at the bottom of the
screen while textures keep v=0 at the top. Frames letterbox via the viewport,
and nextDrawable paces the loop at presentation rate.
Guest-image submission now returns false (callers use their CPU-readback
fallback, which the presenter can show); draw and compute submission still
fail loudly pending later phases.
* [Gpu] Complete the guest-GPU seam: lift the AGC bypass surface onto the backend
The abstraction left AGC and VideoOut calling VulkanVideoPresenter statics
directly for guest work ordering (EnterGuestQueue, SubmitOrderedGuestAction,
SubmitOrderedGuestFlipWait, WaitForGuestWork), guest-image lifecycle (initial
data seeding, writes, fills, extents, upload tracking), the texture-content
cache probe, guest memory attachment, storage-offset alignment, perf
counters, and presenter close. With a non-Vulkan backend selected those
calls silently hit a never-started Vulkan presenter.
All of it now crosses IGuestGpuBackend: the Vulkan backend delegates to the
existing presenter statics (no behavior change), and the Metal backend
answers exactly like a presenter that is not running (sequence 0, image
unknown), which keeps callers on the same inline/CPU fallbacks they take
today. TextureContentIdentity moves to the seam types, and the bounded
AGC-to-presenter transfer pool becomes the backend-neutral GuestDataPool
(one pool by necessity: the AGC layer rents, the presenter returns).
AgcExports snapshots the backend's offset alignment once — it was a const
before and is read in per-draw loops (shader-key hashing, offset rounding).
* [Gpu] Metal guest work queue and guest images: ordered flips, writes, fills, blits
Mirrors the Vulkan presenter's execution model. AGC submissions become work
items consumed by the render loop in logical-guest-queue order: FIFO within
each guest queue, ready queues scheduled round-robin, completion tracked as
a contiguous sequence plus an out-of-order set, and producer backpressure
(count and payload caps) that consumer-enqueued follow-ups bypass to avoid
self-deadlock. The drain is budgeted (12ms, 256 items) so a backlog cannot
starve the Cocoa event pump or the present.
Guest images are Metal textures keyed by guest address, created on first
use from the registered display-buffer format tag (byte-identical encoding
to the Vulkan backend) and seeded once from pending initial data or guest
memory, since PS5 render targets alias guest memory. Coherence mirrors the
Vulkan design: DMA-style writes swap in a freshly written texture (never
mutating one an in-flight present may sample), fills clear through a
hazard-tracked render pass, and same-extent blits copy on the GPU.
Ordered flips capture the named image into an immutable version at their
exact queue position, so later work cannot change the frame a flip
selected; flip waits complete by queue position alone. Presentation picks
the newest ready queued guest frame (retiring superseded captures),
re-resolving mutable address-keyed textures at encode time so a write swap
never leaves a stale handle.
* [Gpu] Metal translated draws: pipelines, render state, bindings, write-back
Executes the seam's translated-draw surface on Metal. Offscreen, depth-only,
and storage draws are ordered guest work rendering into guest-addressed
images (published targets register as flip sources exactly like the Vulkan
backend); onscreen draws and recognized fixed-function draws ride the
presentation and render at present time into a pooled target.
Pipelines are built from GuestRenderState and cached by shader identity plus
a state hash: guest CB blend factor/op codes, write masks (bit-reversed for
MTLColorWriteMask), depth ZFUNC (bit-identical to MTLCompareFunction),
vertex attribute formats decoded from the same guest (dataFormat,
numberFormat) table the Vulkan backend uses, and RDNA 2:10:10:10 mapped to
Metal's R-low-bits 1010102 layout. Guest viewports pass through unchanged —
Metal accepts the negative heights PS5 games program, which is also how the
Vulkan backend inherits its orientation. Rect lists draw as 4-vertex strips;
Metal has no triangle fans, so those degrade to lists with a one-time warn.
Bindings follow the Gen5MslTranslator contract: global buffers at their flat
slot on both stages, SharpEmuUniforms (dispatch limit + buffer byte lengths)
after them, textures/samplers at the image slots with samplers decoded from
the raw guest descriptor words, and vertex streams at slot 26+ so they never
collide. Writable global buffers write back to guest memory before the work
item completes, preserving the CPU-visible GPU-write ordering point that
WaitForGuestWork promises. Feedback reads of a live render target sample a
blit snapshot; pooled guest data returns to GuestDataPool after upload.
Known simplifications for follow-up: textures upload a single mip level, and
the texture-content cache stays unclaimed (IsTextureContentCached=false)
until write-tracker-driven eviction exists, trading upload bandwidth for
correctness.
* [Gpu] Metal compute dispatch: the last seam gap
Guest compute dispatches are ordered guest work like draws. The uniforms
contract carries the per-axis dispatch limit (explicit thread counts when
the guest supplied them, groups x threadgroup size otherwise) so the
kernel's bounds guard clamps the overshoot threads of the last threadgroup;
threadgroup dimensions come from the translated shader, which bakes them at
compile time. Compute pipeline states cache per shader handle.
Storage images are shared live through the guest-image registry: a
dispatch's writes are visible to later draws, blits, and flips of the same
address, the address registers as a flip source at submit, and writer
sequences keep presentation waiting on exactly the work that produced the
frame. Writable buffers write back to guest memory before the work item
completes — the CPU-visible ordering point the returned sequence promises
through WaitForGuestWork.
Metal has no dispatch-base; nonzero base groups execute without the offset
behind a one-time warning until the emitted kernel grows base support.
SHARPEMU_SKIP_ALL_COMPUTE=1 skips all dispatches for hang isolation, same
as the Vulkan backend. With this the Metal backend implements the entire
IGuestGpuBackend surface — nothing throws.
* [Gpu] Address review: real bytes-per-pixel in guest-image uploads
Guest-image uploads hard-coded 4 bytes per texel, which mis-strided
Rgba16*/Rg32Float/Rgba32Float images and, in the guest-memory seed and
storage-snapshot paths, could make replaceRegion read past the managed
buffer. Texel width now comes from the pixel format, and
ReplaceTextureContents clamps the row count to what the source buffer
actually holds, so no caller can overread regardless of pitch and format.
RGBA8 initial data seeds only 4-byte-texel images; wider formats seed from
guest memory, whose layout is the image's native one. Extent byte counts
use the real texel width too.
Also restores the reference's comment on the deliberate single-item
backpressure admit: with no payload outstanding, refusing an oversized item
would wait forever since nothing is left to drain.
* [ShaderCompiler] First real-game fixes: SSendmsg no-op, scalar-state buffer declaration
Bring-up against a real title (2D engine, NGG shaders) found every draw
rejected at translation: RDNA2 NGG shaders bracket their exports with
s_sendmsg (GS_ALLOC_REQ/DEALLOC) to reserve hardware export space, and
neither translator handled the opcode — it fell through to the scalar-ALU
guard and failed with 'missing scalar destination'. Both translators now
treat SSendmsg as a no-op alongside SNop/SWaitcnt: exports are translated
directly, so the hardware message is moot. This was a shared gap, not a
backend one; the Vulkan path would reject the same shaders.
With translation unblocked, the OS Metal compiler rejected the emitted MSL:
the body reads the per-dispatch scalar-state buffer (initial SGPRs plus
per-binding byte biases) as b{initialScalarBufferIndex}, but the kernel
signature only declared the stage's own global bindings, so the name never
existed. The signature now declares it (const device — it is only read) at
its flat slot.
The presenter also logs one line when it first presents real content,
making 'window up but nothing shown' diagnosable from the log alone.
Verified: the title goes from 100% draw misses and a black screen to
~58k translated draws per minute and 4K frames presenting.
* [Gpu] Metal presenter: NSTimer-driven render loop under [NSApp run]
Replaces the hand-pumped event loop with a real running main loop. The
presenter now creates the NSApplication, orders the CAMetalLayer-backed
window on screen, and calls [NSApp run] so Core Animation's run-loop observer
actually commits presented drawables to the window server — without a running
loop the layer never composites and the window stays black regardless of what
is rendered into the drawable.
The per-frame work moves into RenderFrame, driven by a repeating NSTimer on
the main run loop (a tiny NSObject subclass whose onFrame: is an
UnmanagedCallersOnly callback, registered via the ObjC runtime — no binding
package). CADisplayLink is the natural choice and was tried first, but its
callback never fires in this process; proven in isolation against a bare
AppKit harness where a timer fires and composites and the display link does
not — the emulator runs as x86-64 under Rosetta and the display-server-backed
link is not serviced there. nextDrawable still blocks to the display, so the
timer only needs to keep up, not pace precisely.
Also fixes window sizing (the fixed 1280x720 window was being sized from the
guest 4K display mode, which macOS clamps while the layer keeps 4K geometry —
nothing visible), makes the metal layer the view's backing layer (wantsLayer
before setLayer) with an explicit frame, and stops both the AppKit loop and
the CFRunLoop on window close.
* [Gpu] Metal draws: normalize inverted viewports, resolve flips to drawn content
Two correctness fixes surfaced bringing a real title up. Guests program
Vulkan-style negative-height viewports for y-up rendering; Metal's NDC is
already y-up and rasterizes nothing for a negative height, so the viewport is
converted to the equivalent non-inverted rect (origin shifted, height
negated) with the same on-screen mapping.
Ordered flips now prefer produced content. A flip names the display buffer's
start address, but games render into the pixel surface past the buffer's
metadata block; the resolver takes the exact-address image when GPU work
wrote it, else the nearest same-extent GPU-written image within the buffer's
plausible metadata window, else the exact-address image even if only
seeded — so a flip presents the drawn frame rather than an empty seed. A
GpuWritten flag on guest images (set by draws, writes, blits, and dispatch
storage) distinguishes produced content from a speculative guest-memory
seed.
* [ShaderCompiler] Metal translator: per-stage uniforms slot, VCC/EXEC as data, exit branches
Three correctness fixes found by running a real game against the Vulkan
backend's behavior:
- Gen5MslShader carries UniformsBufferIndex: each stage emits its
SharpEmuUniforms argument at globalBufferBase + totalGlobalBufferCount,
and stages sharing a draw can disagree, so the presenter must bind the
buffer per stage (Metal API validation: "missing Buffer binding at
index 7 for sharpemu_uniforms").
- VCC (s106:s107) and EXEC (s126:s127) live in the scalar register file
as raw 32-bit values with the per-lane bools as synced views. Programs
legally park plain data in VCC (s_buffer_load into s[106] and then
v_rcp_f32 of it); the bool-only model returned ballot masks instead.
- A branch to (or past) the program's end is an exit, matching the
SPIR-V translator: sprite alpha-kill shaders use this to skip their
tail and were rejected ("branch target outside program"), silently
dropping every draw that used them.
Also: pixel-stage ballots use the per-lane form (this thread's own bit),
since simd_ballot is undefined inside the divergent dispatcher loop, and
v_readfirstlane returns the lane's own value under that model.
* [Gpu] Metal presenter: per-stage uniforms bind, keyboard input, perf overlay, title parity
- Bind SharpEmuUniforms at each stage's declared slot (see the paired
translator change); one shared index left the vertex stage's slot
unbound, zeroing its bounds-checked loads and killing interpolants.
- Vertex attribute byte offsets move onto the vertex descriptor (buffers
bind at zero) and join the pipeline cache key, which they were silently
missing from once baked into the pipeline.
- Unresolvable draw textures log a throttled warning instead of silently
binding nothing.
- Keyboard input: an NSView subclass records keyDown/keyUp and feeds the
POSIX host-input seam with a Windows-VK to macOS-keycode map covering
the keys pad emulation polls; SHARPEMU_METAL_AUTOKEY scripts key
presses for headless runs.
- Perf overlay (F1) drawn like the Vulkan presenter: CPU-rasterized
panel uploaded to a small texture and composited with the present
pipeline, with RecordPresent/RecordDraw feeding real numbers.
- Window title gains the selected GPU suffix and refreshes when the
guest registers its application name; the layer is marked opaque so
guest alpha never reaches the compositor.
* [Audio] Quiet sceAudioOutOutput on ports disposed by host shutdown
Closing the window disposes audio ports while guest audio threads are
still draining their last buffers; every remaining output then failed
the port lookup and logged a WARN per buffer (~190/s) until process
exit. Report success for missing ports once shutdown has begun; a bad
handle during normal operation still returns INVALID_ARGUMENT.
* [ShaderCompiler] Metal graphics stages model a single logical wave lane
The pixel stage used the real thread_index_in_simdgroup with all-ones
ballots while the vertex stage modeled lane 0 with 1-bit ballots, and
VReadlaneB32 still emitted a real simd_shuffle — reading another
fragment's register. Metal leaves simdgroup ops undefined inside the
divergent while(active){switch(pc)} dispatcher (empirically they
corrupted EXEC reconstruction), so graphics stages cannot use them.
Unify vertex and pixel on the SPIR-V translator's no-subgroup fallback:
one logical wave lane (lane 0), ballots degrade to bit 0, and the
shuffle-select family (readlane, readfirstlane, DPP16/DPP8 selects,
permlane16) resolves to the lane's own value. Writelane keeps the
lane-compare against the constant lane, matching the reference
fallback. Compute is untouched: its threads map one-to-one onto real
simdgroup lanes and still shuffle for real.
Entry parameter lists now always emit trailing commas and are closed by
one helper, so stage-specific trailing parameters no longer dictate
ordering.
* [ShaderCompiler] Metal compute mirrors the SPIR-V translator's wave semantics
Compute threads map one-to-one onto real simdgroup lanes, so restore
real simd_ballot for the compute prelude (the per-lane form was a
graphics fix that swept compute along) — masks parked in VCC/EXEC now
hold each lane's actual bit, and mbcnt/cndmask/saveexec read real
masks. VReadfirstlaneB32 broadcasts from the first guest-active lane
(ballot of EXEC, then ctz), matching the SPIR-V translator's explicit
first-active-lane broadcast rather than SPIR-V BroadcastFirst's
first-host-active semantics.
The wave64 gate moves into the translator and only rejects programs
that contain wave-sensitive operations (the SPIR-V translator's
subgroup-usage predicates: shuffle family, readfirstlane, wave control,
mbcnt, or VCC/EXEC operands). A wave64 kernel without them executes
identically per-thread on 32-wide Apple simdgroups, so it now
translates instead of being dropped.
* [Gpu] Metal draw textures resolve like the Vulkan presenter
Sampling a live guest target previously required the exact current
image at the descriptor's address; anything else silently bound
nothing. Mirror the Vulkan presenter's resolution chain:
- A descriptor naming a guest depth target's write or read address
samples the depth image (identity channel select). Depth32Float
cannot blit to a color format, so the ordered snapshot round-trips
through a private staging buffer into an R32Float texture.
- Replacing a render target at the same guest address (new extent or
format) retires the old image into a bounded variant cache instead of
releasing it, and resolution scores the current image plus variants
by descriptor match — exact extent over view format over
initialization, active image breaking ties. Larger images qualify
only for tiled descriptors, matching IsCompatibleGuestImageAlias.
- The throttled unresolved-texture warning remains the detector for
anything the chain still cannot resolve.
* [ShaderCompiler] Document the Metal translator's wave-size model
Audit outcome for wave64 fidelity, no behavior change: every B64 mask
op, saveexec, and VCCZ/EXECZ test already reads and writes the full
register pair, lane indices never exceed 31 by construction, and the
GPU-executing runtime tests cover 64-bit exec save/restore. Record the
model in the class header.
* [Gpu] Plumb CB_BLEND constant color through both backends
The CONSTANT_COLOR / CONSTANT_ALPHA blend factors were mapped by both
backends but nothing ever supplied the constant, so any draw using them
blended against transparent black. Decode CB_BLEND_RED..ALPHA (the
constants existed unused) into GuestRenderState.BlendConstant, set it
dynamically per draw on both sides: Vulkan declares the blend-constants
dynamic state and calls CmdSetBlendConstants beside the viewport,
Metal calls setBlendColorRed:green:blue:alpha: in EncodeRenderState.
* [Gpu] Metal per-draw uploads bump-allocate from shared arena pages
Every draw created one MTLBuffer and one managed copy per binding
(padded guest globals, uniforms, vertex and index bytes), which
dominated allocation churn at hundreds of MB/s of garbage and held the
guest flip rate well under the display rate. Uploads now bump-allocate
256-aligned slices from 8 MiB shared-storage arena pages and bind by
offset; pages recycle once the last command buffer that referenced
them reports completion, polled at each drain so the ObjC interop
stays block-free.
Write-backs carry the slice's data pointer directly (the page outlives
the command buffer the caller waits on), the alignment-bias contract is
preserved by placing data at the bias inside its slice, and the padded
copies, per-draw buffer releases, and per-draw byte arrays are gone.
In-game on the test title: allocation rate ~795 to ~564 MB/s (the
remainder is the AGC-side per-draw guest snapshots), GC per stats
window ~30/30/17 to 16/16/15, CPU ~150 to ~131%. Metal validation
stays clean.
* [Gpu] Metal draw-texture cache: skip per-draw guest texel copies
Mirror the Vulkan presenter's identity-keyed texture cache: once the
render thread decodes a draw texture, the AGC submit thread skips the
guest-memory read/detile/copy for that identity entirely (the generic
IsTextureContentCached hook, which the Metal backend previously
hardcoded to false) and the render thread serves the cached MTLTexture
without re-uploading. GuestImageWriteTracker write-protects the source
pages; a guest CPU write evicts the entry at the next drain, and the
skip/eviction race self-heals by reading the texels directly.
Eviction differs from Vulkan in one deliberate way: dirty entries are
collected by address rather than identity, since ConsumeDirty clears
the flag on first read and several identities (same texels, different
samplers) can share one address.
Dreaming Sarah in-game on an M5 Max: guest flips 47 -> 60 (display
rate, matching Vulkan), ALLOC 564 -> 41 MB/s, gen0 GC 16 -> 5 per
second, CPU 131% -> 72%. Metal API validation clean; all 25 shader
compiler tests pass.
* [Gpu] Metal snapshot pool: recycle feedback-read textures and staging
Feedback reads created and destroyed an MTLTexture per draw (and for
depth sampling a private staging MTLBuffer too). Pool both with the
upload-arena lifecycle: acquisitions are tagged with the command buffer
that samples them at commit, and return to a bounded free list once it
reports completion. The command queue is serial, so the earlier
snapshot-blit command buffer is necessarily complete by then as well.
Dreaming Sarah renders correctly in-game; Metal API validation clean;
all 25 shader compiler tests pass. (The depth-sample path is exercised
only by inspection — no testable title samples depth yet.)
* [Gpu] Metal batched guest commands: one command buffer per drain
Draws and compute dispatches encode into a shared batch command buffer
committed once per drain instead of one commit per work item, mirroring
the Vulkan presenter's batched guest commands. Ordering inside the
batch is by encoder sequence: draw textures are now pre-resolved before
the consuming render or compute encoder opens, so feedback-read
snapshot blits encode into the batch (after the passes that rendered
the source) rather than committing ahead of them in separate command
buffers. Flips, image writes/blits, ordered actions, CPU-visible
write-backs, and every drain exit flush the batch first, preserving
the serial-queue ordering and WaitForGuestWork contracts.
Dreaming Sarah in-game on an M5 Max: CPU 72% -> 59% at a steady 60
guest flips; Metal API validation clean; all 25 shader compiler tests
pass.
* [Gpu] Metal vertex streams: share buffer slots, reject overflow gracefully
void Terrarium aborts with '-[MTLVertexAttributeDescriptorInternal
setBufferIndex:]: buffer index (31) must be < 31': every vertex
attribute got its own buffer slot from base 26, so six streams walk
past Metal's last vertex-stage buffer index (30) and the framework
assertion kills the process (reported by vladdenisov on PR #283).
Attributes of an interleaved vertex arrive from AGC as one stream
each, all reading the same guest buffer — assign slots by unique
(base address, stride, length) so those share one slot and one
upload. A draw whose unique streams still overflow the range is
skipped with a throttled warning instead of aborting. The assigned
slot keys the pipeline cache alongside the attribute offset, since
aliasing changes the baked vertex descriptor.
Dreaming Sarah renders correctly in-game at 60 flips with Metal API
validation clean; all 25 shader compiler tests pass. (void Terrarium
itself is not testable here — no decrypted copy.)
* [Gpu] Metal: drain guest work on enqueue, not only at render ticks
The Vulkan presenter's render loop is pulsed when guest work arrives
and waits at most a few milliseconds; the Metal render loop drained
guest work only inside its NSTimer tick, so every guest submit-then-
wait round-trip (release-mem labels, event writes, CPU-visible write-
backs) cost up to a full frame interval. Games that chain several such
waits per frame crawl: void Terrarium ran at 14 guest flips against
Vulkan's display rate, and input-to-effect latency suffered everywhere.
Enqueueing guest work now schedules a coalesced onGuestWork: message
onto the main run loop via performSelectorOnMainThread (block-free,
matching the NSTimer trampoline pattern), which drains the queue
immediately. A producer blocked on a full queue schedules the same
wake before waiting. void Terrarium's title menu: 14 -> 59 flips/s;
Dreaming Sarah unchanged at 60 with validation clean.
* [ShaderCompiler] Metal samplers: per-stage compact slots, not texture slots
Sampler argument indices copied the global texture slot (image binding
base + index), but Metal exposes only 16 sampler slots per stage
against 31 texture slots — a draw whose stages sample more than 16
images total emitted [[sampler(16+)]] and the MSL failed to compile
('sampler attribute parameter is out of bounds'), dropping the draw
(void Terrarium's in-game scenes).
Samplers now count sampled (non-storage) images from zero within each
stage, and Gen5MslShader carries the image-index -> sampler-slot map
plus the stage's image binding base so the presenter binds each
stage's samplers exactly where its shader declared them. A stage that
samples more than 16 images fails translation loudly. All 25 shader
compiler tests pass; goldens unchanged (single-texture fixtures keep
sampler 0).
* [Gpu] Metal draw textures: native guest formats, BC blocks, channel select
The draw-texture path assumed every texture was RGBA8: created
Rgba8Unorm, uploaded 4 bytes per pixel, and rejected anything whose
texel copy was smaller than W*H*4 as undersized. Games shipping
BC-compressed atlases (void Terrarium's entire in-game art) rendered
black, and because the rejected textures were never created they were
never content-cached — the AGC layer re-read and re-detiled megabytes
per draw (1.6 GB/s allocation, gen2 collections every second, 8 guest
flips).
Map guest texture formats to Metal case for case with the Vulkan
table (BC1-BC7 upload raw blocks — Mac-family GPUs sample them
natively — plus the 8/16/32-bit linear formats), size expectations
with the same block-aware byte math AGC uses, and honor the
descriptor's DST_SEL channel select through the texture swizzle,
mirroring Vulkan's component mapping. Unmapped codes keep the RGBA8
fallback.
void Terrarium now reaches gameplay past New Game: 49-54 guest
flips (from 8), no undersized-texture warnings, validation clean.
Dreaming Sarah unchanged at 60. All 25 shader compiler tests pass.
* [Gpu] Metal feedback reads: one snapshot per content version
Every draw sampling a live guest image blitted a fresh full-texture
snapshot, so compositing games that sample their render target on
most draws (void Terrarium: ~100 of ~105 draws per frame) pushed
gigabytes per second of blit traffic through the driver.
Guest images now carry a content version, bumped by every draw that
targets them, image write, blit destination, storage dispatch, and
guest-memory seed. The feedback-read path reuses one cached snapshot
until the version moves, so the blit happens per content change
instead of per draw. The image holds the snapshot's retain; consuming
command buffers keep replaced snapshots alive until they complete,
and retire/replace/write paths release the cache with the image.
void Terrarium in-game: 49 -> 58 guest flips at higher draw
throughput (Vulkan reference runs the same scene at 17-20 fps).
Dreaming Sarah unchanged at 60; validation clean; 25/25 tests pass.
* [Core] Pre-visit tracked texture pages before managed guest writes
A managed write into a page the guest-image write tracker has
protected dies with a fatal AccessViolation: the runtime surfaces
SIGSEGV in managed code as an exception before the resumable signal
bridge can restore access, unlike native guest stores which recover
through TryHandleWriteFault. Dead Cells crashed exactly there — an
AGC release-mem label write (CpuContext.TryWriteUInt64 on the render
thread) landing on a page the texture cache tracks.
TryWrite now calls GuestImageWriteTracker.NotifyManagedWrite up
front, unprotecting and dirtying any tracked pages in the span before
the copy — the hook existed for precisely this but had no callers.
Since this puts the tracker on every managed guest-write path, the
range snapshot now carries its overall bounds (one immutable object,
so the intersection test is always consistent with the array), letting
the common no-texture-pages case reject in a few instructions.
Dead Cells no longer crashes; Dreaming Sarah and void Terrarium
unaffected; all 25 shader compiler tests pass.
* [VideoOut] Name the active GPU backend in the macOS window title
macOS can run either backend — Vulkan through MoltenVK or native Metal
via SHARPEMU_GPU_BACKEND — so the window title now ends with the one in
use, e.g. "... · Apple M5 Max (Metal)" or "(Vulkan)". The suffix is
appended in SetSelectedGpuName (the single point both presenters call
to fold in the GPU name) and gated to macOS, so Windows and Linux
titles are unchanged. The name comes from a new BackendName on the
guest-GPU seam.
* [Gpu] Metal window: resizable, native full-screen, live drawable sizing
Add NSWindowStyleMaskResizable so the window can be dragged to any size
and set NSWindowCollectionBehaviorFullScreenPrimary so the green button
enters native full-screen instead of zooming. CAMetalLayer does not
track its drawable size to bounds on its own (even as a view's backing
layer), so the render loop matches drawableSize to the layer's current
bounds x contentsScale before each nextDrawable — a no-op on the common
unchanged tick. The present pass already aspect-fit letterboxes into the
drawable, so any window aspect ratio scales the frame without distortion.
Reading -bounds needs the x86-64 stret ABI for its 32-byte CGRect
return, added as SendStretRect. Verified live: drag-resize and
full-screen both scale correctly with Metal API validation clean.
* [ShaderCompiler] Metal wave64 compute: emulate cross-lane ops via scratch bridge
Replace the wave64 loud rejection with emulation, mirroring the SPIR-V
translator. A 64-lane guest wave is two 32-wide Apple simdgroups
co-resident in one threadgroup (Metal packs thread_index_in_threadgroup
0-31 into simdgroup 0, 32-63 into simdgroup 1), so sharpemu_lane becomes
thread_index_in_threadgroup & 63 and cross-lane ops that span the full
wave rendezvous the two halves through threadgroup scratch:
- ballot into EXEC/VCC/SGPR pairs: each half's simd_ballot is written to
its scratch slot, a threadgroup_barrier syncs, and all lanes recombine
the 64-bit mask into the low/high register pair (centralized in
EmitBallotStore, which the wave32 path shares).
- read-first-lane: broadcasts the lowest active lane's value across both
halves through a scratch slot (EmitWave64ReadFirstLane).
- mbcnt lo/hi: 64-lane thread-mask math (no cross-lane op, just correct
per-lane masks; lanes >= 32 would overflow a 32-bit shift, so split).
The barriers are safe because the guest's scalar PC keeps all 64 lanes
lockstep through the dispatcher. Scope matches the SPIR-V reference: the
scratch is indexed by half, so correct for a one-wave (64-thread)
workgroup, and readlane across halves stays a 32-wide shuffle. Wave-
agnostic wave64 kernels still translate per-thread unchanged.
Verified on the real GPU (MetalRuntimeTests): the emitted wave64 MSL
compiles, and a 64-lane dispatch runs through the bridge barriers
without deadlocking, returning the broadcast value. All 27 tests pass;
Dreaming Sarah (60/60) and void Terrarium (in-game, 58 flips) show the
shared wave32 ballot path is unaffected.
* [Gpu] Metal samplers: bind through an argument buffer, lifting the 16-slot cap
Metal exposes only 16 direct [[sampler(N)]] slots per stage, but real
shaders sample more (void Terrarium's scene shader: 17 images) and were
dropped at translation. Route samplers through a per-stage argument
buffer instead: the MSL declares a Gen5Samplers struct (one sampler per
sampled image, [[id(N)]]) taken as constant& at a buffer slot past the
stage's globals/uniforms/scalar-state, and the runtime writes each
sampler's Tier 2 gpuResourceID into an arena slice bound there. Textures
stay on direct [[texture(N)]] slots (31 is enough). One sampler per
image keeps them distinct, matching the SPIR-V/Vulkan path — no dedup,
so no wrong-sampler artifacts.
Verified argument buffers lift the limit on Apple Silicon (20-sampler
pipeline probe). Dreaming Sarah renders correctly at 60/60 with Metal
API validation clean; the void Terrarium scene shader that exceeded the
limit now compiles and runs (draws 74 -> 102/frame); all 27 shader
compiler tests pass, goldens unchanged (fixtures sample nothing).
* [Gpu] Metal: Shared storage for CPU-populated, GPU-sampled textures
The MTLTextureDescriptor default is Managed, which on unified memory needs
an explicit host->device sync we never issue after replaceRegion, so the
GPU can sample stale texels. These textures are CPU-uploaded and GPU-read,
so Shared (coherent, no sync on Apple Silicon) is the correct mode.
* [HLE] Add missing AGC/AudioOut/Pad exports blocking Unity+FMOD titles
Four exports were unresolved and hard-stalled GPU/audio/input init in
Unity titles (Lunar Lander Beyond froze there before opening VideoOut):
- sceAgcDriverSetTFRing / sceAgcDriverSetHsOffchipParam: tessellation-ring
and hull-shader off-chip config. We translate shaders directly, so these
only need to report success for init to proceed.
- sceAudioOutGetPortState: report a connected primary output at full volume.
- scePadDeviceClassGetExtendedInformation: report a standard pad (no special
peripheral) so device-class probes resolve.
Generic HLE, backend-agnostic (helps the Vulkan path equally).
* [VideoOut] RegisterBuffers2: mask the 32-bit category, accept COMPRESSED
sceVideoOutRegisterBuffers2's category is a 32-bit SceVideoOutBufferCategory
passed on the stack, but we read the full 64-bit slot — whose upper word
carries stale GNM magic (0xC0DEC0DE...) the caller never cleared. The old
check then rejected every call as INVALID_VALUE, so buffer registration
failed and no frame ever presented. Mask to 32 bits and accept both
UNCOMPRESSED (0) and COMPRESSED (1); we present either identically.
Fixes Lunar Lander Beyond reaching its window (now presents 3840x2160).
* [HLE] Stub sceAudioPropagation (3D-audio) so Astro Bot boots past its assert
Astro Bot hard-crashed right after the splash: it calls
sceAudioPropagationSystemQueryMemory during audio init, and because the
whole libSceAudioPropagation module was unimplemented the call failed, so
the game asserted (AudioPropagationContext.cpp:43) and executed int 0x41 to
abort — an unrecoverable trap that kills the process.
We don't model acoustic propagation (geometry-driven reverb/occlusion is a
quality feature, not a correctness gate). The API is placement-style, so
QueryMemory reports a buffer size and the rest succeed as no-ops: the system
lives in the caller's own buffer. All 39 entry points stubbed; the game now
boots past the assert to the presenter. Backend-agnostic HLE.
* [Kernel] pthread_cond_wait: don't spuriously EPERM an untracked mutex
pthread_cond_wait/timedwait required our host-side mutex tracking to show
the calling thread as the owner, else it returned EPERM. But libkernel's
uncontended mutex fast-path locks the mutex word in guest memory directly,
without an HLE call, so we often never observe the lock and see owner==0.
Real pthread_cond_wait requires the caller to hold the mutex but does not
verify it for normal mutexes, so EPERM here is doubly wrong: it spins the
guest (Hades hammered this millions of times/sec) and, worse, skips the
unlock — leaving the mutex held and wedging every thread that later blocks
on pthread_mutex_lock. When the mutex reads as untracked (owner==0), adopt
ownership so the unlock/wait/re-lock cycle is balanced and actually releases
it. Genuine ownership violations (owned by another thread) still error.
Eliminates the EPERM storm and converts the resulting livelock into correct
blocking; no effect on games that lock through the HLE (owner already set).
* [Core] SSE4a EXTRQ patch: read the xmm register from ModRM, not xmm2
The loader rewrites Sony's AMD-only SSE4a EXTRQ+blend idiom into SSE4.1 at
boot, because Rosetta 2 and Intel hosts raise #UD -> SIGILL on EXTRQ. The
matcher hard-coded the source register to xmm2 (ModRM 0xC2), but the compiler
allocates it freely: Dead Cells (PPSA15552) emits the identical idiom against
xmm1, so it slipped through unpatched and the game died with SIGILL right
after the first frame.
Read the register from the ModRM r/m field instead, covering xmm0-xmm7, and
require it to be consistent across the EXTRQ and the blend. The pure
match/encode logic is extracted into Sse4aExtrqBlendPatch, isolated from the
native page-patching, and unit-tested for every register plus the round trip
and rejection cases; DirectExecutionBackend just applies it.
Dead Cells now patches its xmm1 idioms and boots past the first frame.
* [Ngs2] Implement non-allocator sceNgs2SystemCreate / sceNgs2RackCreate
Dead Cells uses the non-allocator NGS2 create entry points, which were
unimplemented. sceNgs2SystemCreate came back as an unresolved import, so the
game got a garbage system handle; every downstream sceNgs2RackCreate /
sceNgs2RackGetVoiceHandle then failed, the voice handle stayed null, and once
gameplay started the audio path polled sceNgs2VoiceGetState/VoiceControl on
the null voice forever — freezing the game in-level at FLIP 0.
The non-allocator forms differ only in a caller buffer (rsi/rcx) vs an
allocator callback; the system/option and out-handle arguments sit at the
same positions, so they alias the existing WithAllocator implementations.
Resolves the NGS2 InvalidVoiceHandle storm (591+/run -> 0).
* [SaveData] Real save subsystem: ~/SharpEmu/Saves/<titleId>, events, full CRUD
Rework the SaveData HLE from a partial stub into a working subsystem:
- Storage moves to ~/SharpEmu/Saves/<titleId>/<dirName>/ (was next to the
exe under user/savedata/<userId>/<titleId>), overridable via
SHARPEMU_SAVEDATA_DIR. Metadata (title/subtitle/detail/userParam) and icon
live in <slot>/sce_sys/. Pure path + param.json logic is isolated in a new
SaveDataStorage type and unit-tested.
- Async event model: sceSaveDataGetEventResult now resolves (was an
unresolved import a save worker polled forever), returning queued completion
events or a clean 'no event' status; SyncSaveDataMemory posts a
SAVE_DATA_MEMORY_SYNC_END event. Plus GetEventInfo/SetEventInfo/register
callbacks.
- New exports: Mount/Mount2/Mount5/Umount, Delete/Delete5, GetParam/SetParam,
SaveIcon/SaveIconByPath/LoadIcon, GetAllSize/GetProgress/GetMountInfo/
IsMounted/GetSaveDataCount/GetMountedSaveDataCount/Abort, Initialize/
Initialize2/Terminate, SaveDataMemory v1 aliases.
- Mounts are tracked so Umount2 really unregisters the /savedata0 mapping
(new KernelMemoryCompatExports.UnregisterGuestPathMount) and params/icons
resolve against the live mount; DirNameSearch surfaces param.json titles.
15 new unit tests (storage layout/sanitize/metadata + mount/event/param/delete
exports); full suite 277 passing.
* [Gpu] Metal: Cmd+F1 toggles Apple's Metal Performance HUD
Plain F1 keeps the built-in CPU-rasterized perf overlay; Cmd+F1 now toggles
the system Metal Performance HUD on the CAMetalLayer, Metal backend only.
Command-modified keys never reach keyDown: (AppKit routes them through the
key-equivalent chain), so the input view gains a performKeyEquivalent:
override that claims Cmd+F1 (also silencing the system beep) and leaves
everything else to the responder chain.
Configured per Apple's 'Customizing Metal Performance HUD':
developerHUDProperties with mode=default + logging=default, plus
MTL_HUD_LOG_SHADER_ENABLED=1 passed directly in the dictionary — HUD,
per-frame statistics logging, and shader-compile logging all enabled from
one property set; mode=disabled hides it again. Guarded by a
respondsToSelector: check for older macOS.
* [Gpu] Metal: also catch Cmd+F1 in keyDown: for the HUD toggle
Function keys reach keyDown: even with Command held (AppKit only reroutes
some chords through performKeyEquivalent:), so the HUD toggle was never
firing there. Handle Cmd+F1 in both the keyDown: and performKeyEquivalent:
paths, and keep it out of MetalHostInput so it can't also flip the plain-F1
perf overlay.
* [Audio] Diagnostics: NGS2 voice-param dump + AudioOut peak-amplitude trace
Two gated traces (idiomatic SHARPEMU_LOG_* style) that pinpoint where audio
dies for NGS2-based games:
- SHARPEMU_LOG_NGS2 now walks the sceNgs2VoiceControl param list and logs each
{size,id} block header + payload bytes, confirming the real layout
(header = u32 size, u32 id; waveform-block param id=0x10000001 carries the
guest PCM pointer at +8; rate param id=0x10000005 carries the resample ratio).
- SHARPEMU_LOG_AUDIO_OUT logs sceAudioOutOutput call count and the peak
amplitude of each submitted buffer.
Finding on void Terrarium: sceAudioOutOutput is called thousands of times on
both 8ch/float32 ports, but every buffer has peak=0.0 — the guest submits pure
silence. The host path (AudioOut -> PCM convert -> CoreAudio) is proven correct;
the silence originates in Ngs2SystemRender, which zeroes the output buffer
instead of mixing voices. Restoring audio for NGS2 games requires a real NGS2
software mixer (next).
* [Audio] NGS2 software mixer: decode + mix PS-ADPCM voices
NGS2-based games were silent because sceNgs2SystemRender only zeroed the
output buffer. This adds a real software mixer:
- Ngs2VagDecoder: clean-room PS-ADPCM ("VAGp") decoder producing mono PCM16
with loop points resolved from the exact per-frame flag values (3=loop
start, 6=loop end, 1/7=one-shot end).
- Voice control now parses the SceNgs2VoiceParamHead command list, decodes the
waveform-blocks param's VAGp container once, and arms the voice.
- sceNgs2SystemRender mixes every armed voice belonging to the system into the
leading grain of the render buffer as interleaved float32 (nearest-sample
resample from the source rate to 48 kHz, additive into the front L/R pair),
which is exactly what games copy to sceAudioOutOutput.
Verified on void Terrarium: previously peak=0.0 silence at AudioOut, now real
audible SFX/music. Voices are still armed on waveform assignment rather than an
explicit kick, so pooled/duplicate voices can overlap — trigger-state handling
is a follow-up.
* [Gpu] AGC: latch GPU-wait satisfaction to the produced value
Fixes a lost-wakeup race that stalled games at a black/splash screen. When a
RELEASE_MEM packet writes a completion label, the guest frequently resets that
label to 0 immediately to reuse it next frame. Our wake path
(GpuWaitRegistry.CollectSatisfied) re-reads *current* guest memory, so if the
reset lands before the wake pass runs, the transient satisfied window is missed
and the suspended DCB waits forever — even though the producing write executed
(traced as wrote=True) and its producer is marked completed.
RELEASE_MEM producers now call GpuWaitRegistry.LatchSatisfiedByValue with the
value they actually wrote, recording satisfaction at the moment of the write for
any waiter that value satisfies. CollectSatisfied honors the latch regardless of
the current (possibly-reset) memory value. This is fail-closed: a waiter only
latches when a real producer wrote a genuinely satisfying value.
Verified: Astro Bot's DEADBEEF sentinel wait (dcb.graphics waiting on a
release_mem label) that was permanently stuck is now resolved; void Terrarium is
unregressed (runs, audio intact, no producerless stalls). Astro still has
separate unresolved blockers (producer-behind-its-own-wait cascades and
producer=none-observed labels) tracked for follow-up; WRITE_DATA/DMA_DATA
producers could latch too but are left out until there is evidence they race.
* [Gpu] AGC: retry indirect dispatches whose GPU-computed dims aren't ready
GPU-driven games (Astro Bot) build their frame on the GPU: a compute dispatch
writes the thread-group dimensions for the next DISPATCH_INDIRECT into a guest
buffer. Our AGC parser reads those dimensions on the CPU at parse time, which
runs before the producing dispatch has executed on the render thread — so it
read 0/0/0 and dropped the work (agc.dispatch_reject zero-dimension), leaving
the scene unrendered (black) and cascading into stuck cross-queue fence waits.
Instead of dropping a zero-dimension INDIRECT dispatch, suspend the DCB on its
dimensions buffer (reusing the WAIT_REG_MEM suspend/resume + GpuWaitRegistry
machinery) until the producer writes non-zero dims, then re-parse and dispatch.
A bounded per-wait deadline (150 ms) resumes-and-drops a genuinely empty
indirect dispatch so it can never stall the queue, making the change
non-regressive: worst case matches the old drop behavior after a short wait.
Direct dispatches (dims inline) are unaffected.
Result: Astro Bot goes from a permanent black screen to actually rendering
(the presenter reports "Metal VideoOut presenting 3840x2160"). void Terrarium —
which issues no indirect dispatches — is unregressed (runs, audio intact, zero
rejects). Astro then hits a separate, newly-reached downstream crash (guest
TBB worker thread_set_state failure) tracked for follow-up.
* [ShaderCompiler] Metal: keep compute shaders within read_write and LDS limits
Two Metal limits made real Astro Bot compute shaders fail to compile/create,
which dropped their dispatches and cascaded into stuck GPU waits (splash hang):
- Textures with access::read_write are capped at 8 per function, but every
storage image was declared read_write. Track each binding's actual access
during body emission (ImageLoad->read, ImageStore->write, ImageAtomic and a
load+store sharing one binding->read_write) and emit the minimal qualifier,
so read-only/write-only storage images no longer count against the cap.
- Threadgroup memory is capped at 32 KB. A shader requesting the full 32 KB of
LDS plus the separate 3-dword wave64 bridge overflowed by 12 bytes. Alias the
bridge into the top of the LDS allocation when both are used, mirroring the
SPIR-V translator's _waveScratchInLds path, keeping the total at 32 KB.
Verified on Astro Bot: "read_write access exceeds maximum (8)" and "Threadgroup
memory size (32780) exceeds maximum (32768)" are both gone; the 27 MSL golden
tests still pass (no golden used a storage image or LDS+wave64 shader).
* [Gpu] AGC: break cross-queue GPU wait deadlocks with a produced-value fallback
Real GPU-driven titles (Astro Bot) drive graphics and compute queues with
mutually dependent WAIT_REG_MEM fences: graphics waits on a compute EOP label,
compute waits on a graphics label. On hardware the two queues run concurrently
so the cycle resolves, but our submission parser is serial, so a label that gets
written -> reset for reuse -> re-waited across queues can wedge forever. The
latch fix helped the write-then-consume race but the cycle re-formed each frame
(graphics stuck at 3 flips, compute queues permanently suspended).
Producers now record the last value they wrote to each label
(GpuWaitRegistry.RecordProduced). A new deadlock breaker
(CollectDeadlockBroken, run from DrainResumableDcbs) releases any waiter stuck
past a 500 ms deadline whose condition is satisfied by that recorded value —
i.e. a real producer signalled the label at least once, guest memory has just
since been reset. It never fabricates a value, and the long deadline means
legitimate fences (which complete within a frame) never trip it.
Verified: Astro Bot goes from 3 flips (wedged on splash) to 25, loads its
splash level ("LevelDocument Loaded: ps_logo") and produces 2432x1368 frame
content. void Terrarium is untouched — 0 deadlock-break events, 1020 flips,
audio intact (its waits resolve far under the deadline). Tunable via
SHARPEMU_GPU_DEADLOCK_BREAK_MS.
* [Cpu] SSE4a EXTRQ patch: cover any blend destination register, not just xmm0
The EXTRQ+VPBLENDD idiom rewrite only matched when the blend destination was
xmm0 (VEX.vvvv byte 0x79). Sony's toolchain allocates that register freely: a
Dead Cells build emits `EXTRQ xmm4,0x28,0x00 ; VPBLENDD xmm3,xmm3,xmm4,2`
(VEX byte 0x61, dest xmm3). That instance stayed unpatched, so the AMD-only
EXTRQ reached Rosetta 2 and raised #UD -> SIGILL (0xC000001D) the moment the
game entered gameplay (loading level PrisonStart).
Read the destination register from VPBLENDD's VEX.vvvv / ModRM.reg as well as
the source from the ModRM r/m field, and emit PINSRD into that destination. Both
are still constrained to xmm0-xmm7 by the fixed VEX prefix. Match/encode stay in
the unit-tested helper.
Verified: Dead Cells now patches 14 EXTRQ blends (previously 0 on this build),
no SIGILL, and reaches PrisonStart. 15 patch unit tests pass, including the exact
xmm3/xmm4 bytes that faulted.
* [HLE] Implement Dead Cells' remaining unresolved imports
Three imports Dead Cells calls during boot/level-load were unresolved, so they
returned no defined value:
- scePadGetHandle (libScePad): returns the primary pad's handle (polled every
frame for input); same validation as scePadOpen.
- sceNpEntitlementAccessGetAddcontEntitlementInfo (libSceNpEntitlementAccess):
singular add-on-content lookup; we own no DLC, so zero the info out and return
OK, matching the existing list variant.
- sceNpUniversalDataSystemEventPropertyArraySetString: telemetry setter, dropped.
Dead Cells now boots with zero unresolved imports. (It still stalls later at
PrisonStart level-load — a separate GPU/threading issue, not an import gap.)
* [Kernel] Fix pthread mutex deadlock: trylock semantics + stale-waiter clog
Hades hard-froze during boot on a "free but reserved" mutex: owner==0 yet
every acquisition failed forever. Two independent defects in the pthread
mutex compat layer combined to wedge it, both traced from real runs.
1. trylock incorrectly required an empty wait queue. POSIX
pthread_mutex_trylock succeeds whenever the mutex is not currently held
and owes no fairness to queued waiters; gating it on Waiters.Count==0 made
a spin-on-trylock loop (which the game runs) spin forever against a single
undrainable waiter even though owner==0. trylock now acquires on owner==0;
the blocking lock still honours FIFO so genuine blocked waiters are not
starved by a barging locker.
2. cond_timedwait timeouts leaked mutex re-acquire waiters. A cond wait's
timeout enqueues a re-acquire waiter whose wake hand-off can be lost,
orphaning it in the mutex queue. Multiple orphans from one thread piled at
the FIFO head; the unlock hand-off then woke a dead wake-key and the mutex
never drained. A thread can hold at most one pending acquisition on a
mutex, so EnqueueMutexWaiterLocked now prunes any prior waiter for the same
thread before enqueueing — collapsing the leaked pile.
Verified: Hades advances from a hard freeze at ~4.9M HLE calls (main and a
worker both blocked on the same free mutex) to 24.9M calls with no stall,
reaching the save-data/user-service boot stage. void tRrLM behaves
identically with and without the change (no regression); all 268 Libs tests
pass.
|
||
|
|
6dda6589d0 |
test: add Kernel/Loader unit tests (22 tests) (#373)
- SelfLoader: reject unknown magic, truncated headers; parse PS5 SELF embedded ELF - KernelMemory: MapNamedFlexibleMemory/mprotect/munmap argument validation - KernelEventQueue: create/delete/add/trigger/wait lifecycle Co-authored-by: OMP <omp@local> |
||
|
|
a709ccca17 | [shader_recompiler] Fix guest image byte count calculation for Vulkan video presenter (#395) | ||
|
|
5309f384cf | Reject undefined numeric LogLevel values (#390) | ||
|
|
f84d869795 |
[Kernel] Match path cache comparisons to host filesystem case sensitivity (#381)
The negative-stat cache and the apr file-size cache memoize host
filesystem probe outcomes, but both were keyed with an ignore-case
comparer while the probes themselves (File.Exists/Directory.Exists/
FileInfo) are case-sensitive on Linux. That aliases distinct paths:
- stat("/app0/DATA.BIN") fails, the miss is cached, and a later
stat("/app0/Data.bin") is answered NOT_FOUND from the cache without
ever probing the disk - even though the file exists and the probe
would succeed.
- sceKernelAprResolveFilepathsToIdsAndFileSizes serves the cached size
of a case-distinct sibling file instead of the file's own size.
The registered-mount containment guard had the inverse problem: the
ignore-case StartsWith accepted a ".." path that resolves into a
sibling directory differing from the mount root only by case
("…/Save" vs "…/save"), letting guest I/O escape the mount.
All three sites now compare with the host filesystem's semantics:
ordinal-ignore-case on Windows, ordinal elsewhere. Windows behavior is
unchanged. Tests probe actual host filesystem behavior with real temp
files and skip their case-specific sections on case-insensitive hosts.
|
||
|
|
13269797bf | Add live debugger frontend and mutex stall recovery (#383) | ||
|
|
1c8cdd6537 |
[VideoOut] Initialize output options storage (#315)
* [VideoOut] Initialize output options storage * [VideoOut] Keep output options size local |
||
|
|
41c9b44a8a | [AJM] Track registered codec instance lifecycle (#352) | ||
|
|
743fe5cc26 |
[ShaderCompiler/Vulkan] Match vertex input numeric types (#351)
Declare UINT and SINT vertex attributes with integer SPIR-V component types so shader interfaces match the Vulkan pipeline formats. Keep normalized, scaled, and floating-point formats on float inputs. Signed-off-by: missatjuhvdk1 <177474143+missatjuhvdk1@users.noreply.github.com> Co-authored-by: missatjuhvdk1 <177474143+missatjuhvdk1@users.noreply.github.com> |
||
|
|
0b1dea43e8 | [CPU] Preserve blocked leaf import waiters (#350) | ||
|
|
b479dc0466 |
videoout: implement output support query (#269)
Co-authored-by: Chris Cheng <chris@appxtream.com> |
||
|
|
22bbb4e909 |
[AvPlayer] Resolve guest media within app0 (#347)
Handle project-relative file URIs through the guest app0 mount, including unambiguous case-insensitive lookup for case-sensitive hosts. Reject host paths, traversal underflow, malformed or remote URIs, and symlink/reparse escapes; cover accepted app0 forms and sandbox boundaries with nonparallel tests. |
||
|
|
3c500d2cf0 |
[SystemService] Write notice skip flag as byte (#346)
The Gen5 caller supplies a one-byte flag. Preserve pointer and memory-fault behavior while writing only that byte, and cover a seeded guest-memory boundary that rejects the former four-byte write. |
||
|
|
bcb0ebd991 |
[ShaderCompiler] Fix VReadlane scalar destination field (#344)
V_READLANE uses the gfx10 VOP3A vdst byte even though its result is scalar. Decode bits 0-7 and cover the public LLVM s5 and s101 encodings so the VOP3B sdst field cannot be confused with this opcode again. |
||
|
|
ecbb0db9be |
[Kernel] Preserve socket descriptors after failed connect (#343)
Keep ownership of a socket descriptor with the guest when connect fails, and route generic close calls through the socket table. Add a deterministic regression test for the failure and close sequence. Signed-off-by: missatjuhvdk1 <177474143+missatjuhvdk1@users.noreply.github.com> Co-authored-by: missatjuhvdk1 <177474143+missatjuhvdk1@users.noreply.github.com> |
||
|
|
cc290f860b | [Kernel] Return largest available direct-memory span (#334) | ||
|
|
aa25f6e978 |
[Loader] Restore PS5 SELF header support (#342)
Accept both PS4 and PS5 SELF signatures without treating version, key type, or flags as layout markers. Add synthetic coverage for valid header variants and malformed structural fields. Signed-off-by: missatjuhvdk1 <177474143+missatjuhvdk1@users.noreply.github.com> Co-authored-by: missatjuhvdk1 <177474143+missatjuhvdk1@users.noreply.github.com> |
||
|
|
b9eee497aa |
[VideoOut] Scope the AMD integrated-GPU penalty to Windows (#339)
* [VideoOut] Prefer real integrated GPUs over software rasterizers (#325) Penalize only AMD integrated GPUs (the #97 vkCreateGraphicsPipelines crash) instead of all integrated devices, so Intel/Apple/Qualcomm iGPUs outrank Cpu-type software rasterizers (Mesa lavapipe). Hoist ScorePhysicalDevice to the outer class and add unit tests for the ordering. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * [VideoOut] Scope AMD iGPU penalty to Windows via a penalty helper Extract ComputeDevicePenalty (the value subtracted from a device's base score) and gate the #97 AMD-integrated penalty on Windows only. Mesa RADV on Linux (e.g. the Steam Deck's AMD APU) is a different, working driver and should keep its full integrated score. Add a Steam Deck test case and drop inline comments. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * [VideoOut] Move device-scoring helpers out of the const block Relocate ScorePhysicalDevice and ComputeDevicePenalty below the leading const cluster instead of splitting it, and trim the vendor-ID reference comment to adapters an x86-64 host can realistically enumerate. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a9a4366ef4 |
test: cover Gen5 decoder fail-closed boundaries (#336)
Reimplement the five synthetic regressions from codefusion-repo/sharpemu#3 and #8 on sharpemu/sharpemu main
|
||
|
|
faf3a397a8 |
[VideoOut] Fall back to 8-bit RGBA for unknown pixel formats instead of silently failing (#294)
Changes MapPixelFormatToGuestTextureFormat to default to format 56 (8-bit RGBA) when the game uses a pixel format not yet in the known list, with a stderr warning that reports the exact format value for project issue reports. Previously, unknown formats returned 0, which caused RegisterKnownDisplayBuffer to skip registration entirely. The GPU backend then couldn't find the buffer during flip, producing vk.flip_capture_failed, and some games later hit a Debug.Assert in ExecuteOrderedGuestFlipWait. The fallback produces wrong colors for the affected games but lets them render and display output, which is strictly better than a black screen or access violation crash. The pixel format is printed to stderr so developers can identify it and add proper support. Co-authored-by: meowman <haadii2005@gamil.com> |
||
|
|
488b285ecb |
[SaveData] Implement save data memory2 exports (#297)
* [SaveData] Implement save data memory2 exports Astro Bot calls sceSaveDataSetupSaveDataMemory2 during boot and asserts and null-writes when it fails, so the missing import surfaces as a named crash. This implements setup plus the companion get, set, and sync operations that make it useful. The store is one zero-filled file per user and title at sce_sdmemory/memory.dat under the save root, readiness is the backing file's existence, and get, set, and sync return MEMORY_NOT_READY before setup. Struct offsets follow the publicly documented homebrew savedata headers. * [SaveData] Write setup result before mutating the memory backing file |
||
|
|
1a4a2902d4 |
[HLE] Stub ContentExport, Font, and Pad calls Astro Bot needs to boot (#298)
Astro Bot asserts and null-writes when a subsystem init call fails, so each missing import surfaces as a named crash. This stubs the blockers observed during bring-up: content export init, eight font calls, and pad tilt correction. sceFontGetHorizontalLayout writes the same invented geometry as sceFontGetRenderCharGlyphMetrics and the rest report success. Together with save data memory2 these take the title to its splash image and font glyph rendering path. |
||
|
|
6b1abc1a38 | perf(shader): cache pixel export masks (#288) | ||
|
|
7494792249 |
[Tests] Add VirtualMemory and PhysicalVirtualMemory edge case tests (#266)
VirtualMemory: 4 new tests covering unmapped address returns false, zero-length operations, page-boundary access, and cross-gap failures. PhysicalVirtualMemory: 5 new tests covering lazy commit on demand, reserve-only GetPointer, unmapped GetPointer returns null, free-list first-fit reuse, and coalescing both neighbours on middle-range free. Total: 9 new tests, 18 passed (8 VirtualMemory + 5 PhysicalVirtualMemory + 5 GuestMemoryAllocator). No production code changes. |
||
|
|
3585519007 | [Audio] Correct float PCM endpoint conversion (#291) | ||
|
|
dabf723b3e | [Vulkan] Honor guest depth clear state (#290) | ||
|
|
2db1fae282 | [Tests/HLE] Cover APR resolve, stat, and streaming flow (#272) | ||
|
|
33f96252da | [CPU] Reject context transfers to unmapped guest addresses (#273) | ||
|
|
f7981a7ed7 | [Tests/CPU] Verify import trampoline volatile-state ABI (#274) | ||
|
|
e10efa3ae1 | [HLE] Make guest printf formatting locale-invariant (#271) | ||
|
|
16a2131b67 | [PlayGo] Reject authoritative unknown chunk loci (#242) | ||
|
|
9883a9445d |
[AGC] Implement RDNA2 buffer/image/DS 32-bit atomic instructions (#222)
Adds decode and SPIR-V translation for the missing MUBUF, MIMG and DS atomic instructions in the Gen5 shader translator, generalizing the existing BufferAtomicAdd path. Covers swap, cmpswap, add, sub, smin/smax, umin/umax, and/or/xor, inc and dec, plus the DS RTN variants. Image atomics go through OpImageTexelPointer on the storage image binding. Notable: DS_CMPST operand order (DATA0 = comparator, DATA1 = new value) is reversed relative to buffer/image cmpswap, which a dedicated test locks in. ATOMIC_INC/DEC are approximated with OpAtomicIIncrement/IDecrement, exact for the common 0xFFFFFFFF clamp. Verified with 9 new synthetic decoder and end-to-end SPIR-V tests (part of the #36 test corpus effort); full suite passes 35/35. Signed-off-by: missatjuhvdk1 <177474143+missatjuhvdk1@users.noreply.github.com> Co-authored-by: missatjuhvdk1 <177474143+missatjuhvdk1@users.noreply.github.com> |
||
|
|
f2d9051358 | Fix macOS fixed-address allocation collisions (#246) | ||
|
|
53d414e096 |
Bound Vulkan host buffer pool memory (#201)
Co-authored-by: Dafenx <196083014+Dafenxz0@users.noreply.github.com> |
||
|
|
52d2874fa8 |
cpu: emulate BMI1/BMI2/ABM instructions in software when the host lacks them (#249)
## What The native backend runs guest code directly on the host CPU. When the host doesn't implement a BMI1/BMI2/ABM instruction that the PS5's Zen 2 cores do, it raises #UD (STATUS_ILLEGAL_INSTRUCTION). Today the vectored handler just logs the faulting bytes and gives up, so the title dies. This adds a software fallback. On an illegal-instruction fault we decode the opcode with Iced (the decoder already used elsewhere in the backend), evaluate it against the trapped register/memory state, write the result and flags back into the CONTEXT record, step RIP past the instruction, and resume. Instructions covered (32- and 64-bit): ANDN, BLSI, BLSMSK, BLSR, BEXTR, BZHI, TZCNT, LZCNT, RORX, SARX, SHLX, SHRX, PDEP, PEXT. ## Why Users on CPUs without these extensions currently can't get past code that uses them. This is a generic fix (no game-specific hacks) that improves compatibility on older hosts. MULX is intentionally left out for now — its dest_hi/dest_lo operand ordering is easy to get subtly wrong, so I'd rather add it separately with its own tests. ## How it's structured - `BmiInstructionEmulator` holds the pure bit/flag semantics with no dependency on the unsafe CONTEXT plumbing, so it can be unit-tested directly. - `DirectExecutionBackend.IllegalInstruction.cs` is the thin unsafe adapter (decode → read operands → emulate → write back → advance RIP). Anything it doesn't fully model returns false and falls through to the existing diagnostics unchanged, so it can never mis-handle an opcode it doesn't recognize. - One hook in `DirectExecutionBackend.Exceptions.cs`, next to the other TryRecover* calls. - Emits a single one-time "emulating in software" log line, not per-instruction spam. ## How I verified - Added xUnit tests covering every instruction in both widths plus the CF/ZF/SF/OF edge cases (src == 0, shift-count masking, index beyond operand width, etc.). - Cross-checked all the expected values against an independent reference implementation written from the Intel/AMD definitions; results match. - `dotnet build` + `dotnet test` pass locally. ## Notes New files follow .editorconfig (4-space, SPDX headers, REUSE-compliant). |
||
|
|
a8be1daca4 |
Fix sceNetHtonl/Htons/Ntohl/Ntohs discarding their converted result (#211)
Each of the four byte-order helpers computed the swap into Rax and then returned via ctx.SetReturn(0). SetReturn writes its argument into Rax, so it immediately overwrote the converted value with 0 and the guest saw every sceNetHtonl / sceNetHtons / sceNetNtohl / sceNetNtohs call return 0. Leave the swapped value in Rax and return ORBIS_GEN2_OK as the dispatch status instead, matching how the value-returning exports in this file (e.g. sceNetPoolCreate) already work. Add a NetExports test suite covering the swaps, the 16-bit width masking, an htonl/ntohl round-trip, and a regression guard that a non-zero input never converts to 0. |
||
|
|
26dbad8ac8 | fix(input): map POSIX stick endpoints exactly (#244) | ||
|
|
b4b95014f1 |
Fix/deadcells crash (#262)
* Boot compatibility fixes for UE titles, GUI toggles, and DeS render/boot work Checkpoint of the Monster Truck Championship and Demon's Souls boot work. Each piece is independently useful and verified against the titles. - playgo: scePlayGoGetLocus now returns BAD_CHUNK_ID for chunk ids outside the known set, matching real firmware. Titles enumerate chunk ids until that error; answering OK for every id made the scan wrap the ushort range and spin forever. Missing-sidecar and no-app0 fallbacks report a fully-installed single chunk 0 so scePlayGoOpen keeps succeeding. - kernel: restore the SHARPEMU_WRITABLE_APP0 opt-in. Unpackaged UE dumps write their Saved tree under /app0 during PS5 component init and treat the denial as a fatal boot error. - pad: accept handle 0 as the primary pad across all pad calls. Real firmware hands out small non-negative handles and some titles read state with handle 0. - bthid: env-gated experiment hooks for the Thrustmaster wheel middleware investigation (fail-only-RegisterCallback modes and a synthetic enumeration callback with a zeroed event struct). All default off. - gui: add SHARPEMU_LOG_IO and SHARPEMU_WRITABLE_APP0 toggles to the Environment tab. - videoout: per-swapchain-image render-finished semaphores (the shared semaphore raced the swapchain); whole-mip-chain layout init for offscreen guest images (sampled binds read mips stuck in Undefined); GPU-resident texture availability now canonicalizes through the texture format table and accepts compatibility-class aliases, cutting per-frame CPU texture re-reads (143 GB -> 55 GB per 300 s in Demon's Souls, 0.2 -> 0.5 fps). - hle: add sceSystemServiceGetNoticeScreenSkipFlag, sceSystemServiceGetMainAppTitleId (title id published from the runtime), and sceNpWebApi2CreateUserContext (refuses so the online layer backs off). - rtc: SHARPEMU_RTC_PROBE_RANGE diagnostic dumps the code around a busy-wait caller of sceRtcGetCurrentTick once; costs nothing when unset. * [ShaderCompiler] Fix VReadlaneB32 scalar destination field The scalar destination lives in the low vdst byte (bits 0-7); it was read from bits 8-14, the VOP3B carry-out field readlane does not have, sending every readlane result to s0. Verified against raw gfx10 encodings and LLVM's assembler tests (v_readlane_b32 s5, v1, s2 -> low byte 0x05). * [VideoOut] Survive device loss and flip-order asserts without dying Two ways a frame could take down the whole presenter: - Device loss between any two Vulkan calls in a frame unwound the window thread, and the Dispose-time fence check then threw again, masking the original error. Catch the loss at the frame boundary, retire presentations and guest submissions whose fences can never signal, and keep the window loop pumping so the game (audio, logic) carries on. - The ordered-flip capture invariant is violated ~100 times per run by Demon's Souls (PPSA01342); on debug builds the Debug.Assert fail-fasts the process with nothing in the log. Downgrade it to a once-per-version warning until the capture/wait ordering is understood. * Fix Dead Cells shader cache regression --------- Co-authored-by: StealUrKill <35749471+StealUrKill@users.noreply.github.com> |
||
|
|
9bacb883f1 |
Fix Linux aligned mapping retention (#247)
* Fix Linux aligned mapping retention * Cover Linux aligned mapping retention |
||
|
|
864cbb0fa0 |
[AGC/Vulkan] Extend PS5 runtime and rendering compatibility (#216)
* [Core] Add POSIX native execution and PS5 SELF support Extend the native backend, guest TLS, fixed-address memory, and loader paths needed by PS5 titles on Windows, Linux, and macOS. Keep workstation GC so high-core-count hosts do not reserve over fixed guest image bases. * [HLE] Expand PS5 service and media compatibility Add the kernel, threading, save-data, networking, audio, video-codec, font, dialog, and service exports required by newer PS5 software. Preserve every SysAbi NID currently registered by main while adding the compatibility surface used by ASTRO BOT. * [AGC/Vulkan] Extend Gen5 shader and presentation support Expand PM4 handling, Gen5 shader translation, MRT and packed export support, guest image tracking, depth initialization, texture aliasing, and Vulkan presentation. Add the performance overlay and address-filtered diagnostics used to validate ASTRO BOT with original shaders. * [Core] Align static TLS reservation across hosts * [Pad] Align primary user ID with UserService * [Gpu] Preserve runtime scalar buffers across renderer seam * [AGC] Restore omitted command helper exports * [Vulkan] Reuse primary views for promoted MRT targets * [Vulkan] Preserve scratch storage bindings in compute dispatches |
||
|
|
f69fdd4027 | Fix PNG chunk CRC validation (#241) | ||
|
|
1be009ce40 |
[HLE] Fix POSIX condition variable semantics (#113) (#223)
* [HLE] Trigger AGC graphics events by filter instead of exact ident (#173) The PM4 EVENT_WRITE packet carries a 6-bit hardware EVENT_TYPE, but the guest registers AGC events via sceAgcDriverAddEqEvent with a full guest eventId. These two values are not the same numbering scheme, so the exact ident lookup in TriggerRegisteredEvents never matched and the AGC interrupt thread hung forever. Add TriggerRegisteredEventsByFilter, which wakes every graphics event registration on every queue. This is a compatibility workaround for issue #173 while the real PS5 mapping remains unknown. Includes unit tests covering the mismatched ident/eventType case. * [HLE] Fix POSIX condition variable semantics (#113) Remove PendingSignals from PthreadCondState. POSIX condition signals are edges, not semaphore credits - a signal with no waiter must have no effect. The previous implementation persisted signals, causing lock inversions and predicate bypasses. Changes: - Remove PendingSignals property and TryConsumePendingSignal method - Remove pending signal consumption logic from PthreadCondWaitCore - Remove PendingSignals increment from PthreadCondSignalCore - Add regression tests verifying POSIX-correct behavior Fixes #113 |