Commit Graph

78 Commits

Author SHA1 Message Date
Slick Daddy 33be88bdf9 memory: back the free pages of a partially-overlapping fixed mapping (#458)
A SCE_KERNEL_MAP_FIXED request whose window partially overlaps an
existing allocation was failing outright: AllocateAt reserves the whole
range in one all-or-nothing VirtualAlloc, which returns 0 on partial
overlap. The mapping call then returned NOT_FOUND while leaving the free
tail unmapped, so the guest faulted (0xC0000005) writing into it.

Add IGuestAddressSpace.TryBackFixedRange, which walks the range via the
host Query (VirtualQuery reports contiguous same-state runs) and fills
only the free sub-ranges, leaving already-backed pages untouched. This
matches the fixed-mapping contract on hardware. Route the fixed
reservation path through it via a new backPartialOverlap flag.

Co-authored-by: slick-daddy <slick-daddy@users.noreply.github.com>
2026-07-20 09:08:29 +03:00
Spooks 90c72ebecf Fixes a Mutex Issue Preventing Some UE Titles From Booting (#451)
* Optimize guest import, memory, and pthread hot paths

* Fix UE adaptive mutex self-lock handling
2026-07-19 13:20:05 -06:00
Nekono 8ef5a54ee4 cpu: emulate AMD-only Zen 2 instructions in software (#449)
Handle immediate EXTRQ and INSERTQ as well as MONITORX and MWAITX when the host raises illegal-instruction faults. Add unit coverage for SSE4a bit-field semantics and preserve existing load-time patching.

Co-authored-by: zocomputer <help@zocomputer.com>
2026-07-19 21:57:42 +03:00
StealUrKill bc51cc2c4d Prevent invalid SaveData writes from damaging guest memory (#444)
Add an optional write monitor so the team can find future memory damage on each supported desktop system.
2026-07-19 20:27:17 +03:00
wearr a60bfc9c83 [Kernel] Implement pthread semaphore exports (#424) 2026-07-19 04:18:16 +03:00
Berk a030cb5a5d Gpu runtime stalls (#410)
* [runtime] restore default GC mode

* [cpu] add string leaf stubs

* [ampr] allow concurrent reads

* [bink] keep guest decode path

* [kernel] streamline host memory access

* [shader] add scalar memory fallback

* [gpu] bound guest data pool

* [gpu] reduce queue stalls

* [video] stabilize guest resources

* revert lock file
2026-07-19 00:31:50 +03:00
Dafenx 336286e588 CPU: scan final TLS access pattern offset (#414)
Co-authored-by: Dafenx <196083014+Dafenxz0@users.noreply.github.com>
2026-07-19 00:07:32 +03:00
Spooks daaeb6213e Fix Massive Bug Preventing UE5 Titles From Booting (#406)
* Fix cross platform memcpy bug
2026-07-18 12:50:59 -06:00
Gutemberg Ribeiro 94153955b0 [Gpu] Metal backend: complete IGuestGpuBackend implementation on AppKit + Metal (#283)
* [ShaderCompiler.Metal] MSL translator core: dispatcher, EXEC model, compute stage

The Metal codegen backend, rebuilt on the merged backend-neutral
abstractions (replacing the pre-abstraction spike): consumes
(Gen5ShaderState, Gen5ShaderEvaluation) and emits MSL text; the renderer
owns MTLLibrary compilation, mirroring the emitters-produce-bytes rule
the Vulkan sibling documents.

The execution model mirrors Gen5SpirvTranslator: one invocation per GCN
lane (wave32 — natively the Apple simdgroup width), a typeless uint
register file with as_type<float> bitcasts, EXEC/VCC as per-lane bools
whose guest-visible mask registers materialize via simd_ballot, and the
same PC-dispatcher loop over basic blocks with the
SHARPEMU_SHADER_MAX_STEPS iteration guard and the dominating-scalar-
definition dataflow for buffer binding resolution. Unlike SPIR-V, MSL
permits shared prelude functions, so unaligned/subdword buffer access
is a range-checked device-uchar* helper instead of per-site inlining;
buffer byte lengths and the compute dispatch limit travel in one
reserved SharpEmuUniforms constant buffer (Metal has no OpArrayLength).

This first slice covers the compute entry point end to end: scalar/
vector ALU core (moves, int/float arithmetic, FMA family, shifts,
bitfield ops, min/max/med3, conversions, transcendentals with the Tau
scale on sin/cos), the full VCmp/VCmpx compare matrix writing VCC/EXEC,
the saveexec family, scalar compares and SCC-updating SOP2 forms, lane
ops (readfirstlane, mbcnt), VOP3 abs/neg/clamp/omod modifiers, scalar
memory, and raw global/buffer loads, stores, and atomics with EXEC
guards. Unsupported opcodes fail loudly with pc + mnemonic. SDWA/DPP,
typed format loads, LDS, images, and the pixel/vertex stages follow in
the next phases.

Tests live in their own self-contained project (the per-backend model:
depends only on the codegen under test): hand-assembled synthetic
fixtures drive the real decoder end to end, structural assertions and
golden-MSL comparisons run on every platform since translation is pure
text generation, and goldens regenerate via SHARPEMU_UPDATE_GOLDENS=1.

* [ShaderCompiler.Metal] Real-device runtime tests: compile + execute on the GPU

Lifts the spike's objc_msgSend LibraryImport harness (MTLDevice /
MTLCompileOptions with fast-math off, as a real Metal backend must
compile) onto the new translator contract: the guest data buffer binds
at index 0 and the SharpEmuUniforms constant buffer (dispatch limit +
buffer byte lengths) at index 1.

Three runtime tiers on hosts with a Metal device (no-op elsewhere so
Windows/Linux CI stays green): every fixture's emitted MSL must be
accepted by the OS runtime Metal compiler; the exec-store program must
produce bit-exact GPU results including the EXEC-masked store that must
not land; and a scalar countdown loop must iterate through the PC
dispatcher (s_cmp_lg_u32 + s_cbranch_scc1 across five round trips).
The loop fixture also fixes its own hand-assembly: s_sub_i32 sets SCC
to signed overflow, not result-nonzero, so the loop condition uses an
explicit compare.

* [ShaderCompiler.Metal] Phase 2: scalar/vector ALU parity with the SPIR-V translator

Ports the remaining ALU semantics from Gen5SpirvTranslator.Alu so the two
codegens cannot disagree on instruction behavior:

- Carry/borrow family (v_add_co/_ci, v_sub_co/_rev, v_subb/_rev) with the
  carry mask written to the VOP3 scalar destination or VCC, ANDed with
  EXEC; v_mad_u64_u32 with the 64-bit pair result and carry-out.
- Full SDWA support: byte/word source selects with sign-extension,
  integer abs/neg modifiers, and destination-select merge (zero-fill,
  sign-extend, preserve) into the previous register value.
- DPP16/DPP8: quad permute, row shl/shr/ror, mirror/half-mirror,
  broadcast and xor controls via simd_shuffle, bound-control and
  row/bank write-enable masks, fetch-inactive handling; DPP-predicated
  compares merge into VCC.
- Lane ops: readfirstlane from the first EXEC-active lane via
  ballot+ctz, readlane/writelane, permlane16/permlanex16.
- 64-bit scalar ops over SGPR pairs in real ulong arithmetic (logic
  family, shifts, bfe/bfm with width clamping, wqm quad expansion,
  cselect, mov, s_getpc) plus the B32/B64 saveexec families.
- Sopk forms decode the signed 16-bit immediate and s_cmpk compares the
  destination register; SOPC scalar compares including s_bitcmp0/1.
- Conversions: f16<->f32 via as_type<half>, pkrtz with round-to-zero
  mantissa truncation, pknorm via pack_float_to_{s,u}norm2x16, pk_u8
  byte insert, off_f32_i4 table, rpi/flr rounding; cube id/sc/tc/ma
  decision trees; v_cmp_class_f32; VCCZ/EXECZ/SCC readable as data.

Also fixes a real phase-1 bug the reference surfaced: fmamk/fmaak
sources arrive in natural order from the decoder, so all MAD/FMA forms
are fma(src0, src1, src2) — the previous operand swap computed
v1*v2+K for v_fmamk (should be v1*K+v2). The regenerated fmac golden
shows the corrected expansion, and mirroring the SPIR-V translator,
v_mul_u32_u24 is a full 32-bit multiply (only the hi/mad forms mask).

All 13 Metal tests pass including the real-GPU execution tier.

* [ShaderCompiler.Metal] Phase 3: typed format loads, LDS, and D16 subdword memory

Typed MUBUF/MTBUF loads convert through the descriptor's GFX10 unified
format at execution time, mirroring the SPIR-V translator: the prelude
bakes a 128-entry format table from the shared Gfx10UnifiedFormat
decoder (compiled shaders may be reused with new SRDs, so decoding must
stay dynamic), per-component layouts for the legacy DATA_FORMAT values
drive range-checked unaligned loads, NUM_FORMAT conversion handles
unorm/snorm (clamped at -1)/uscaled/sscaled/uint/sint/float including
f16 and the 10/11-bit unsigned mini-floats of 10_11_11 / 11_11_10, the
missing-component default is one in the format's domain, and dst_sel
swizzling comes from descriptor word 3. Format stores stay raw dword
stores like the reference.

LDS lands as 32 KB of threadgroup memory (gated on the program actually
using DS ops so occupancy is not taxed): ds_read/write b32/b64/b96/b128,
the write2/read2 pairs including st64 scaling, and ds_add_u32 as a
relaxed threadgroup atomic, with EXEC-guarded writes and the address
masked into bounds. Subdword loads/stores gain the D16/D16Hi variants
that merge into one half of the destination register (and shift the
source for high stores), classified the same way as the reference.

New GPU-executed fixture: an LDS round trip (write literal, s_barrier,
read back, store to the buffer) passes bit-exact on a real Metal device
alongside the existing tiers.

* [ShaderCompiler.Metal] Phase 4: pixel stage, images, and interpolation

The pixel entry points land with the same contract as the SPIR-V
translator (single-target and MRT forms, validated for unique guest
slots and dense host locations): the emitted fragment function takes a
stage_in struct carrying [[position]] plus the interpolated attributes
discovered from the program's V_INTERP controls, writes an output
struct with one [[color(hostLocation)]] attachment per binding typed by
its Float/Sint/Uint kind, seeds pixel-input VGPRs in SPI_PS_INPUT_ADDR
compact order from the fragment coordinate, keeps EXEC masking through
translation, and discards lanes that exit with EXEC off. Exports write
MRT targets per component under EXEC (disabled components keep their
previous value) including compressed half-pair exports; vertex-target
exports no-op until the vertex stage.

Images arrive as texture2d<float|int|uint> arguments (storage bindings
as access::read_write) with samplers alongside, classified from the
descriptor's unified format via the shared Gfx10UnifiedFormat decoder
and resolved per instruction with the same dominating-scalar-definition
scheme as buffers. The sample matrix covers implicit LOD, SampleL/Lz,
SampleB, SampleD gradients, PCF compare (manual reference<=texel,
broadcast r,r,r,1), per-lane texel offsets folded into normalized
coordinates by the selected mip extent (Metal sample offsets must be
constants), gather4 including compare and offset forms, clamped
ImageLoad/Mip, bounds-checked EXEC-guarded ImageStore, GetResinfo, and
A16 packed addresses / D16 packed data in both directions.

Graphics stages model LDS as per-invocation scratch (the SPIR-V
Private-array trick) instead of threadgroup memory. Wave ops keep the
invocation's real simdgroup in every stage — Apple fragment simdgroups
make that the same model as compute, where the SPIR-V translator
instead emulates a single logical lane; both round-trip EXEC masks
consistently.

The pixel fixture (interpolated attr0.xy plus inline constants exported
to MRT0) is golden-pinned, structurally asserted, and accepted by the
OS Metal compiler on a real device.

* [ShaderCompiler.Metal] Phase 5: vertex stage and fixed presenter shaders

The vertex entry point completes the four-entry-point contract: the
emitted vertex function takes fetched attributes as a stage_in struct
([[attribute(location)]], bound by the backend via MTLVertexDescriptor
from the reflected vertex inputs), returns [[position]] plus one
[[user(locnN)]] param output per export target 32..63 — unioned with
requiredVertexOutputCount so Metal's exact vertex-out/fragment-in
interface match succeeds, with unexported locations zero-filled —
seeds v5/v8 from [[vertex_id]]/[[instance_id]], intercepts buffer loads
the evaluator captured as fixed-function vertex inputs, and applies the
same EXEC-selected component rules to position/param exports (disabled
components default to 0,0,0,1) including compressed half pairs.

MSL vertex functions have no simdgroup attributes, so the vertex stage
models a single logical wave lane exactly like the SPIR-V translator's
graphics path: lane 0, ballot degrades to 0/1, and lane-shuffle ops
would fail Metal compilation loudly (no real guest vertex shader uses
them).

MslFixedShaders mirrors SpirvFixedShaders for the presenter surface:
the fullscreen-triangle vertex stage (position from the vertex index,
screen-space UV broadcast to every requested attribute location), the
copy/solid/attribute diagnostic fragments, and the output-free
depth-only fragment. Metal forbids "main", so each carries a stable
entry name.

The vertex fixture (constant position + one param export) is
golden-pinned and structurally asserted; it and all five fixed shaders
are accepted by the OS Metal compiler on a real device.

* [ShaderCompiler.Metal] Author static MSL blocks as template files

The prelude helpers (buffer access, ballot, tables), the format-load
conversion functions, and all five fixed presenter shaders move out of
AppendLine walls into Templates/*.msl embedded resources — real Metal
source with syntax highlighting and reviewable diffs — rendered by a
small {{placeholder}} substituter that fails loudly on any
unsubstituted token. Substitution points are deliberately few: the
stage-dependent ballot expression (vertex has no simdgroup attributes),
the baked GFX10 format table and layout cases, and the fixed shaders'
parameters. Per-instruction body emission stays programmatic, where a
template cannot express it.

Behavior-identical by construction: the golden files are untouched and
the whole suite — including the real-device execution and compile
tiers over the templated output — passes against them unchanged. The
.msl files carry no license headers (they would leak into every emitted
shader), so REUSE.toml annotates the Templates directory instead.

* [ShaderCompiler.Metal] Cover the MSL goldens in REUSE.toml

The golden files are verbatim emitter output regenerated by the test
suite; license headers inside them would either break the byte-exact
comparison or force the emitter to write SPDX text into every shader.
Annotate the Goldens directory like the Templates one.

* [ShaderCompiler.Metal] Address review: dominating-binding parity and harness binding indices

The buffer-binding fallback now mirrors the SPIR-V translator: a candidate
binding is accepted only when the descriptor registers hold the exact same
scalar definitions at the target PC as at one of the binding's own access
points (HasSameScalarDefinitions), instead of merely being non-conflicting
at the target. Resolutions are cached per PC like the reference.

The runtime test harness no longer hardcodes buffer indices 0/1 and a
20-byte uniforms blob: TryExecuteSingleThread takes the data/uniforms bind
indices, and ExecuteOrThrow derives them plus the uniforms size from the
compiled shader's GlobalMemoryBindings per the translator contract.

* [Gpu] Add the Metal guest-GPU backend: shader compilation and formats

First phase of the Metal backend behind the IGuestGpuBackend seam. The
backend compiles all three shader stages through Gen5MslTranslator and
exposes the guest render-target format table (mirroring the Vulkan table
case for case; guest format 9 maps to BGR10A2, the Metal layout matching
Vulkan's A2R10G10B10 pack). Wave64 compute is rejected with a clear error
until the two-pass emulation exists.

SHARPEMU_GPU_BACKEND=metal opts in on macOS; Vulkan stays the default on
every platform until the Metal presenter reaches parity. Presenter-side
methods fail loudly instead of dropping guest frames silently.

* [Gpu] Add the Metal presenter core: AppKit window, CAMetalLayer, CPU-frame path

The presenter opens an NSWindow hosting a CAMetalLayer and drives a manually
pumped NSApplication event loop, structured like the Vulkan presenter's
poll-and-render loop and posted onto HostMainThread the same way (AppKit
traps off the process main thread). All OS access goes through objc_msgSend
LibraryImport bindings declared locally — no windowing or binding packages on
this path, which is what keeps it NativeAOT-clean. Struct-returning ObjC
calls are avoided entirely so one calling convention works under Rosetta.

Presents CPU-produced BGRA frames and the splash through a fullscreen
triangle with a dedicated present fragment stage that flips V: with Metal's
y-up NDC the shared fullscreen triangle puts UV (0,0) at the bottom of the
screen while textures keep v=0 at the top. Frames letterbox via the viewport,
and nextDrawable paces the loop at presentation rate.

Guest-image submission now returns false (callers use their CPU-readback
fallback, which the presenter can show); draw and compute submission still
fail loudly pending later phases.

* [Gpu] Complete the guest-GPU seam: lift the AGC bypass surface onto the backend

The abstraction left AGC and VideoOut calling VulkanVideoPresenter statics
directly for guest work ordering (EnterGuestQueue, SubmitOrderedGuestAction,
SubmitOrderedGuestFlipWait, WaitForGuestWork), guest-image lifecycle (initial
data seeding, writes, fills, extents, upload tracking), the texture-content
cache probe, guest memory attachment, storage-offset alignment, perf
counters, and presenter close. With a non-Vulkan backend selected those
calls silently hit a never-started Vulkan presenter.

All of it now crosses IGuestGpuBackend: the Vulkan backend delegates to the
existing presenter statics (no behavior change), and the Metal backend
answers exactly like a presenter that is not running (sequence 0, image
unknown), which keeps callers on the same inline/CPU fallbacks they take
today. TextureContentIdentity moves to the seam types, and the bounded
AGC-to-presenter transfer pool becomes the backend-neutral GuestDataPool
(one pool by necessity: the AGC layer rents, the presenter returns).

AgcExports snapshots the backend's offset alignment once — it was a const
before and is read in per-draw loops (shader-key hashing, offset rounding).

* [Gpu] Metal guest work queue and guest images: ordered flips, writes, fills, blits

Mirrors the Vulkan presenter's execution model. AGC submissions become work
items consumed by the render loop in logical-guest-queue order: FIFO within
each guest queue, ready queues scheduled round-robin, completion tracked as
a contiguous sequence plus an out-of-order set, and producer backpressure
(count and payload caps) that consumer-enqueued follow-ups bypass to avoid
self-deadlock. The drain is budgeted (12ms, 256 items) so a backlog cannot
starve the Cocoa event pump or the present.

Guest images are Metal textures keyed by guest address, created on first
use from the registered display-buffer format tag (byte-identical encoding
to the Vulkan backend) and seeded once from pending initial data or guest
memory, since PS5 render targets alias guest memory. Coherence mirrors the
Vulkan design: DMA-style writes swap in a freshly written texture (never
mutating one an in-flight present may sample), fills clear through a
hazard-tracked render pass, and same-extent blits copy on the GPU.

Ordered flips capture the named image into an immutable version at their
exact queue position, so later work cannot change the frame a flip
selected; flip waits complete by queue position alone. Presentation picks
the newest ready queued guest frame (retiring superseded captures),
re-resolving mutable address-keyed textures at encode time so a write swap
never leaves a stale handle.

* [Gpu] Metal translated draws: pipelines, render state, bindings, write-back

Executes the seam's translated-draw surface on Metal. Offscreen, depth-only,
and storage draws are ordered guest work rendering into guest-addressed
images (published targets register as flip sources exactly like the Vulkan
backend); onscreen draws and recognized fixed-function draws ride the
presentation and render at present time into a pooled target.

Pipelines are built from GuestRenderState and cached by shader identity plus
a state hash: guest CB blend factor/op codes, write masks (bit-reversed for
MTLColorWriteMask), depth ZFUNC (bit-identical to MTLCompareFunction),
vertex attribute formats decoded from the same guest (dataFormat,
numberFormat) table the Vulkan backend uses, and RDNA 2:10:10:10 mapped to
Metal's R-low-bits 1010102 layout. Guest viewports pass through unchanged —
Metal accepts the negative heights PS5 games program, which is also how the
Vulkan backend inherits its orientation. Rect lists draw as 4-vertex strips;
Metal has no triangle fans, so those degrade to lists with a one-time warn.

Bindings follow the Gen5MslTranslator contract: global buffers at their flat
slot on both stages, SharpEmuUniforms (dispatch limit + buffer byte lengths)
after them, textures/samplers at the image slots with samplers decoded from
the raw guest descriptor words, and vertex streams at slot 26+ so they never
collide. Writable global buffers write back to guest memory before the work
item completes, preserving the CPU-visible GPU-write ordering point that
WaitForGuestWork promises. Feedback reads of a live render target sample a
blit snapshot; pooled guest data returns to GuestDataPool after upload.

Known simplifications for follow-up: textures upload a single mip level, and
the texture-content cache stays unclaimed (IsTextureContentCached=false)
until write-tracker-driven eviction exists, trading upload bandwidth for
correctness.

* [Gpu] Metal compute dispatch: the last seam gap

Guest compute dispatches are ordered guest work like draws. The uniforms
contract carries the per-axis dispatch limit (explicit thread counts when
the guest supplied them, groups x threadgroup size otherwise) so the
kernel's bounds guard clamps the overshoot threads of the last threadgroup;
threadgroup dimensions come from the translated shader, which bakes them at
compile time. Compute pipeline states cache per shader handle.

Storage images are shared live through the guest-image registry: a
dispatch's writes are visible to later draws, blits, and flips of the same
address, the address registers as a flip source at submit, and writer
sequences keep presentation waiting on exactly the work that produced the
frame. Writable buffers write back to guest memory before the work item
completes — the CPU-visible ordering point the returned sequence promises
through WaitForGuestWork.

Metal has no dispatch-base; nonzero base groups execute without the offset
behind a one-time warning until the emitted kernel grows base support.
SHARPEMU_SKIP_ALL_COMPUTE=1 skips all dispatches for hang isolation, same
as the Vulkan backend. With this the Metal backend implements the entire
IGuestGpuBackend surface — nothing throws.

* [Gpu] Address review: real bytes-per-pixel in guest-image uploads

Guest-image uploads hard-coded 4 bytes per texel, which mis-strided
Rgba16*/Rg32Float/Rgba32Float images and, in the guest-memory seed and
storage-snapshot paths, could make replaceRegion read past the managed
buffer. Texel width now comes from the pixel format, and
ReplaceTextureContents clamps the row count to what the source buffer
actually holds, so no caller can overread regardless of pitch and format.
RGBA8 initial data seeds only 4-byte-texel images; wider formats seed from
guest memory, whose layout is the image's native one. Extent byte counts
use the real texel width too.

Also restores the reference's comment on the deliberate single-item
backpressure admit: with no payload outstanding, refusing an oversized item
would wait forever since nothing is left to drain.

* [ShaderCompiler] First real-game fixes: SSendmsg no-op, scalar-state buffer declaration

Bring-up against a real title (2D engine, NGG shaders) found every draw
rejected at translation: RDNA2 NGG shaders bracket their exports with
s_sendmsg (GS_ALLOC_REQ/DEALLOC) to reserve hardware export space, and
neither translator handled the opcode — it fell through to the scalar-ALU
guard and failed with 'missing scalar destination'. Both translators now
treat SSendmsg as a no-op alongside SNop/SWaitcnt: exports are translated
directly, so the hardware message is moot. This was a shared gap, not a
backend one; the Vulkan path would reject the same shaders.

With translation unblocked, the OS Metal compiler rejected the emitted MSL:
the body reads the per-dispatch scalar-state buffer (initial SGPRs plus
per-binding byte biases) as b{initialScalarBufferIndex}, but the kernel
signature only declared the stage's own global bindings, so the name never
existed. The signature now declares it (const device — it is only read) at
its flat slot.

The presenter also logs one line when it first presents real content,
making 'window up but nothing shown' diagnosable from the log alone.
Verified: the title goes from 100% draw misses and a black screen to
~58k translated draws per minute and 4K frames presenting.

* [Gpu] Metal presenter: NSTimer-driven render loop under [NSApp run]

Replaces the hand-pumped event loop with a real running main loop. The
presenter now creates the NSApplication, orders the CAMetalLayer-backed
window on screen, and calls [NSApp run] so Core Animation's run-loop observer
actually commits presented drawables to the window server — without a running
loop the layer never composites and the window stays black regardless of what
is rendered into the drawable.

The per-frame work moves into RenderFrame, driven by a repeating NSTimer on
the main run loop (a tiny NSObject subclass whose onFrame: is an
UnmanagedCallersOnly callback, registered via the ObjC runtime — no binding
package). CADisplayLink is the natural choice and was tried first, but its
callback never fires in this process; proven in isolation against a bare
AppKit harness where a timer fires and composites and the display link does
not — the emulator runs as x86-64 under Rosetta and the display-server-backed
link is not serviced there. nextDrawable still blocks to the display, so the
timer only needs to keep up, not pace precisely.

Also fixes window sizing (the fixed 1280x720 window was being sized from the
guest 4K display mode, which macOS clamps while the layer keeps 4K geometry —
nothing visible), makes the metal layer the view's backing layer (wantsLayer
before setLayer) with an explicit frame, and stops both the AppKit loop and
the CFRunLoop on window close.

* [Gpu] Metal draws: normalize inverted viewports, resolve flips to drawn content

Two correctness fixes surfaced bringing a real title up. Guests program
Vulkan-style negative-height viewports for y-up rendering; Metal's NDC is
already y-up and rasterizes nothing for a negative height, so the viewport is
converted to the equivalent non-inverted rect (origin shifted, height
negated) with the same on-screen mapping.

Ordered flips now prefer produced content. A flip names the display buffer's
start address, but games render into the pixel surface past the buffer's
metadata block; the resolver takes the exact-address image when GPU work
wrote it, else the nearest same-extent GPU-written image within the buffer's
plausible metadata window, else the exact-address image even if only
seeded — so a flip presents the drawn frame rather than an empty seed. A
GpuWritten flag on guest images (set by draws, writes, blits, and dispatch
storage) distinguishes produced content from a speculative guest-memory
seed.

* [ShaderCompiler] Metal translator: per-stage uniforms slot, VCC/EXEC as data, exit branches

Three correctness fixes found by running a real game against the Vulkan
backend's behavior:

- Gen5MslShader carries UniformsBufferIndex: each stage emits its
  SharpEmuUniforms argument at globalBufferBase + totalGlobalBufferCount,
  and stages sharing a draw can disagree, so the presenter must bind the
  buffer per stage (Metal API validation: "missing Buffer binding at
  index 7 for sharpemu_uniforms").
- VCC (s106:s107) and EXEC (s126:s127) live in the scalar register file
  as raw 32-bit values with the per-lane bools as synced views. Programs
  legally park plain data in VCC (s_buffer_load into s[106] and then
  v_rcp_f32 of it); the bool-only model returned ballot masks instead.
- A branch to (or past) the program's end is an exit, matching the
  SPIR-V translator: sprite alpha-kill shaders use this to skip their
  tail and were rejected ("branch target outside program"), silently
  dropping every draw that used them.

Also: pixel-stage ballots use the per-lane form (this thread's own bit),
since simd_ballot is undefined inside the divergent dispatcher loop, and
v_readfirstlane returns the lane's own value under that model.

* [Gpu] Metal presenter: per-stage uniforms bind, keyboard input, perf overlay, title parity

- Bind SharpEmuUniforms at each stage's declared slot (see the paired
  translator change); one shared index left the vertex stage's slot
  unbound, zeroing its bounds-checked loads and killing interpolants.
- Vertex attribute byte offsets move onto the vertex descriptor (buffers
  bind at zero) and join the pipeline cache key, which they were silently
  missing from once baked into the pipeline.
- Unresolvable draw textures log a throttled warning instead of silently
  binding nothing.
- Keyboard input: an NSView subclass records keyDown/keyUp and feeds the
  POSIX host-input seam with a Windows-VK to macOS-keycode map covering
  the keys pad emulation polls; SHARPEMU_METAL_AUTOKEY scripts key
  presses for headless runs.
- Perf overlay (F1) drawn like the Vulkan presenter: CPU-rasterized
  panel uploaded to a small texture and composited with the present
  pipeline, with RecordPresent/RecordDraw feeding real numbers.
- Window title gains the selected GPU suffix and refreshes when the
  guest registers its application name; the layer is marked opaque so
  guest alpha never reaches the compositor.

* [Audio] Quiet sceAudioOutOutput on ports disposed by host shutdown

Closing the window disposes audio ports while guest audio threads are
still draining their last buffers; every remaining output then failed
the port lookup and logged a WARN per buffer (~190/s) until process
exit. Report success for missing ports once shutdown has begun; a bad
handle during normal operation still returns INVALID_ARGUMENT.

* [ShaderCompiler] Metal graphics stages model a single logical wave lane

The pixel stage used the real thread_index_in_simdgroup with all-ones
ballots while the vertex stage modeled lane 0 with 1-bit ballots, and
VReadlaneB32 still emitted a real simd_shuffle — reading another
fragment's register. Metal leaves simdgroup ops undefined inside the
divergent while(active){switch(pc)} dispatcher (empirically they
corrupted EXEC reconstruction), so graphics stages cannot use them.

Unify vertex and pixel on the SPIR-V translator's no-subgroup fallback:
one logical wave lane (lane 0), ballots degrade to bit 0, and the
shuffle-select family (readlane, readfirstlane, DPP16/DPP8 selects,
permlane16) resolves to the lane's own value. Writelane keeps the
lane-compare against the constant lane, matching the reference
fallback. Compute is untouched: its threads map one-to-one onto real
simdgroup lanes and still shuffle for real.

Entry parameter lists now always emit trailing commas and are closed by
one helper, so stage-specific trailing parameters no longer dictate
ordering.

* [ShaderCompiler] Metal compute mirrors the SPIR-V translator's wave semantics

Compute threads map one-to-one onto real simdgroup lanes, so restore
real simd_ballot for the compute prelude (the per-lane form was a
graphics fix that swept compute along) — masks parked in VCC/EXEC now
hold each lane's actual bit, and mbcnt/cndmask/saveexec read real
masks. VReadfirstlaneB32 broadcasts from the first guest-active lane
(ballot of EXEC, then ctz), matching the SPIR-V translator's explicit
first-active-lane broadcast rather than SPIR-V BroadcastFirst's
first-host-active semantics.

The wave64 gate moves into the translator and only rejects programs
that contain wave-sensitive operations (the SPIR-V translator's
subgroup-usage predicates: shuffle family, readfirstlane, wave control,
mbcnt, or VCC/EXEC operands). A wave64 kernel without them executes
identically per-thread on 32-wide Apple simdgroups, so it now
translates instead of being dropped.

* [Gpu] Metal draw textures resolve like the Vulkan presenter

Sampling a live guest target previously required the exact current
image at the descriptor's address; anything else silently bound
nothing. Mirror the Vulkan presenter's resolution chain:

- A descriptor naming a guest depth target's write or read address
  samples the depth image (identity channel select). Depth32Float
  cannot blit to a color format, so the ordered snapshot round-trips
  through a private staging buffer into an R32Float texture.
- Replacing a render target at the same guest address (new extent or
  format) retires the old image into a bounded variant cache instead of
  releasing it, and resolution scores the current image plus variants
  by descriptor match — exact extent over view format over
  initialization, active image breaking ties. Larger images qualify
  only for tiled descriptors, matching IsCompatibleGuestImageAlias.
- The throttled unresolved-texture warning remains the detector for
  anything the chain still cannot resolve.

* [ShaderCompiler] Document the Metal translator's wave-size model

Audit outcome for wave64 fidelity, no behavior change: every B64 mask
op, saveexec, and VCCZ/EXECZ test already reads and writes the full
register pair, lane indices never exceed 31 by construction, and the
GPU-executing runtime tests cover 64-bit exec save/restore. Record the
model in the class header.

* [Gpu] Plumb CB_BLEND constant color through both backends

The CONSTANT_COLOR / CONSTANT_ALPHA blend factors were mapped by both
backends but nothing ever supplied the constant, so any draw using them
blended against transparent black. Decode CB_BLEND_RED..ALPHA (the
constants existed unused) into GuestRenderState.BlendConstant, set it
dynamically per draw on both sides: Vulkan declares the blend-constants
dynamic state and calls CmdSetBlendConstants beside the viewport,
Metal calls setBlendColorRed:green:blue:alpha: in EncodeRenderState.

* [Gpu] Metal per-draw uploads bump-allocate from shared arena pages

Every draw created one MTLBuffer and one managed copy per binding
(padded guest globals, uniforms, vertex and index bytes), which
dominated allocation churn at hundreds of MB/s of garbage and held the
guest flip rate well under the display rate. Uploads now bump-allocate
256-aligned slices from 8 MiB shared-storage arena pages and bind by
offset; pages recycle once the last command buffer that referenced
them reports completion, polled at each drain so the ObjC interop
stays block-free.

Write-backs carry the slice's data pointer directly (the page outlives
the command buffer the caller waits on), the alignment-bias contract is
preserved by placing data at the bias inside its slice, and the padded
copies, per-draw buffer releases, and per-draw byte arrays are gone.

In-game on the test title: allocation rate ~795 to ~564 MB/s (the
remainder is the AGC-side per-draw guest snapshots), GC per stats
window ~30/30/17 to 16/16/15, CPU ~150 to ~131%. Metal validation
stays clean.

* [Gpu] Metal draw-texture cache: skip per-draw guest texel copies

Mirror the Vulkan presenter's identity-keyed texture cache: once the
render thread decodes a draw texture, the AGC submit thread skips the
guest-memory read/detile/copy for that identity entirely (the generic
IsTextureContentCached hook, which the Metal backend previously
hardcoded to false) and the render thread serves the cached MTLTexture
without re-uploading. GuestImageWriteTracker write-protects the source
pages; a guest CPU write evicts the entry at the next drain, and the
skip/eviction race self-heals by reading the texels directly.

Eviction differs from Vulkan in one deliberate way: dirty entries are
collected by address rather than identity, since ConsumeDirty clears
the flag on first read and several identities (same texels, different
samplers) can share one address.

Dreaming Sarah in-game on an M5 Max: guest flips 47 -> 60 (display
rate, matching Vulkan), ALLOC 564 -> 41 MB/s, gen0 GC 16 -> 5 per
second, CPU 131% -> 72%. Metal API validation clean; all 25 shader
compiler tests pass.

* [Gpu] Metal snapshot pool: recycle feedback-read textures and staging

Feedback reads created and destroyed an MTLTexture per draw (and for
depth sampling a private staging MTLBuffer too). Pool both with the
upload-arena lifecycle: acquisitions are tagged with the command buffer
that samples them at commit, and return to a bounded free list once it
reports completion. The command queue is serial, so the earlier
snapshot-blit command buffer is necessarily complete by then as well.

Dreaming Sarah renders correctly in-game; Metal API validation clean;
all 25 shader compiler tests pass. (The depth-sample path is exercised
only by inspection — no testable title samples depth yet.)

* [Gpu] Metal batched guest commands: one command buffer per drain

Draws and compute dispatches encode into a shared batch command buffer
committed once per drain instead of one commit per work item, mirroring
the Vulkan presenter's batched guest commands. Ordering inside the
batch is by encoder sequence: draw textures are now pre-resolved before
the consuming render or compute encoder opens, so feedback-read
snapshot blits encode into the batch (after the passes that rendered
the source) rather than committing ahead of them in separate command
buffers. Flips, image writes/blits, ordered actions, CPU-visible
write-backs, and every drain exit flush the batch first, preserving
the serial-queue ordering and WaitForGuestWork contracts.

Dreaming Sarah in-game on an M5 Max: CPU 72% -> 59% at a steady 60
guest flips; Metal API validation clean; all 25 shader compiler tests
pass.

* [Gpu] Metal vertex streams: share buffer slots, reject overflow gracefully

void Terrarium aborts with '-[MTLVertexAttributeDescriptorInternal
setBufferIndex:]: buffer index (31) must be < 31': every vertex
attribute got its own buffer slot from base 26, so six streams walk
past Metal's last vertex-stage buffer index (30) and the framework
assertion kills the process (reported by vladdenisov on PR #283).

Attributes of an interleaved vertex arrive from AGC as one stream
each, all reading the same guest buffer — assign slots by unique
(base address, stride, length) so those share one slot and one
upload. A draw whose unique streams still overflow the range is
skipped with a throttled warning instead of aborting. The assigned
slot keys the pipeline cache alongside the attribute offset, since
aliasing changes the baked vertex descriptor.

Dreaming Sarah renders correctly in-game at 60 flips with Metal API
validation clean; all 25 shader compiler tests pass. (void Terrarium
itself is not testable here — no decrypted copy.)

* [Gpu] Metal: drain guest work on enqueue, not only at render ticks

The Vulkan presenter's render loop is pulsed when guest work arrives
and waits at most a few milliseconds; the Metal render loop drained
guest work only inside its NSTimer tick, so every guest submit-then-
wait round-trip (release-mem labels, event writes, CPU-visible write-
backs) cost up to a full frame interval. Games that chain several such
waits per frame crawl: void Terrarium ran at 14 guest flips against
Vulkan's display rate, and input-to-effect latency suffered everywhere.

Enqueueing guest work now schedules a coalesced onGuestWork: message
onto the main run loop via performSelectorOnMainThread (block-free,
matching the NSTimer trampoline pattern), which drains the queue
immediately. A producer blocked on a full queue schedules the same
wake before waiting. void Terrarium's title menu: 14 -> 59 flips/s;
Dreaming Sarah unchanged at 60 with validation clean.

* [ShaderCompiler] Metal samplers: per-stage compact slots, not texture slots

Sampler argument indices copied the global texture slot (image binding
base + index), but Metal exposes only 16 sampler slots per stage
against 31 texture slots — a draw whose stages sample more than 16
images total emitted [[sampler(16+)]] and the MSL failed to compile
('sampler attribute parameter is out of bounds'), dropping the draw
(void Terrarium's in-game scenes).

Samplers now count sampled (non-storage) images from zero within each
stage, and Gen5MslShader carries the image-index -> sampler-slot map
plus the stage's image binding base so the presenter binds each
stage's samplers exactly where its shader declared them. A stage that
samples more than 16 images fails translation loudly. All 25 shader
compiler tests pass; goldens unchanged (single-texture fixtures keep
sampler 0).

* [Gpu] Metal draw textures: native guest formats, BC blocks, channel select

The draw-texture path assumed every texture was RGBA8: created
Rgba8Unorm, uploaded 4 bytes per pixel, and rejected anything whose
texel copy was smaller than W*H*4 as undersized. Games shipping
BC-compressed atlases (void Terrarium's entire in-game art) rendered
black, and because the rejected textures were never created they were
never content-cached — the AGC layer re-read and re-detiled megabytes
per draw (1.6 GB/s allocation, gen2 collections every second, 8 guest
flips).

Map guest texture formats to Metal case for case with the Vulkan
table (BC1-BC7 upload raw blocks — Mac-family GPUs sample them
natively — plus the 8/16/32-bit linear formats), size expectations
with the same block-aware byte math AGC uses, and honor the
descriptor's DST_SEL channel select through the texture swizzle,
mirroring Vulkan's component mapping. Unmapped codes keep the RGBA8
fallback.

void Terrarium now reaches gameplay past New Game: 49-54 guest
flips (from 8), no undersized-texture warnings, validation clean.
Dreaming Sarah unchanged at 60. All 25 shader compiler tests pass.

* [Gpu] Metal feedback reads: one snapshot per content version

Every draw sampling a live guest image blitted a fresh full-texture
snapshot, so compositing games that sample their render target on
most draws (void Terrarium: ~100 of ~105 draws per frame) pushed
gigabytes per second of blit traffic through the driver.

Guest images now carry a content version, bumped by every draw that
targets them, image write, blit destination, storage dispatch, and
guest-memory seed. The feedback-read path reuses one cached snapshot
until the version moves, so the blit happens per content change
instead of per draw. The image holds the snapshot's retain; consuming
command buffers keep replaced snapshots alive until they complete,
and retire/replace/write paths release the cache with the image.

void Terrarium in-game: 49 -> 58 guest flips at higher draw
throughput (Vulkan reference runs the same scene at 17-20 fps).
Dreaming Sarah unchanged at 60; validation clean; 25/25 tests pass.

* [Core] Pre-visit tracked texture pages before managed guest writes

A managed write into a page the guest-image write tracker has
protected dies with a fatal AccessViolation: the runtime surfaces
SIGSEGV in managed code as an exception before the resumable signal
bridge can restore access, unlike native guest stores which recover
through TryHandleWriteFault. Dead Cells crashed exactly there — an
AGC release-mem label write (CpuContext.TryWriteUInt64 on the render
thread) landing on a page the texture cache tracks.

TryWrite now calls GuestImageWriteTracker.NotifyManagedWrite up
front, unprotecting and dirtying any tracked pages in the span before
the copy — the hook existed for precisely this but had no callers.
Since this puts the tracker on every managed guest-write path, the
range snapshot now carries its overall bounds (one immutable object,
so the intersection test is always consistent with the array), letting
the common no-texture-pages case reject in a few instructions.

Dead Cells no longer crashes; Dreaming Sarah and void Terrarium
unaffected; all 25 shader compiler tests pass.

* [VideoOut] Name the active GPU backend in the macOS window title

macOS can run either backend — Vulkan through MoltenVK or native Metal
via SHARPEMU_GPU_BACKEND — so the window title now ends with the one in
use, e.g. "... · Apple M5 Max (Metal)" or "(Vulkan)". The suffix is
appended in SetSelectedGpuName (the single point both presenters call
to fold in the GPU name) and gated to macOS, so Windows and Linux
titles are unchanged. The name comes from a new BackendName on the
guest-GPU seam.

* [Gpu] Metal window: resizable, native full-screen, live drawable sizing

Add NSWindowStyleMaskResizable so the window can be dragged to any size
and set NSWindowCollectionBehaviorFullScreenPrimary so the green button
enters native full-screen instead of zooming. CAMetalLayer does not
track its drawable size to bounds on its own (even as a view's backing
layer), so the render loop matches drawableSize to the layer's current
bounds x contentsScale before each nextDrawable — a no-op on the common
unchanged tick. The present pass already aspect-fit letterboxes into the
drawable, so any window aspect ratio scales the frame without distortion.

Reading -bounds needs the x86-64 stret ABI for its 32-byte CGRect
return, added as SendStretRect. Verified live: drag-resize and
full-screen both scale correctly with Metal API validation clean.

* [ShaderCompiler] Metal wave64 compute: emulate cross-lane ops via scratch bridge

Replace the wave64 loud rejection with emulation, mirroring the SPIR-V
translator. A 64-lane guest wave is two 32-wide Apple simdgroups
co-resident in one threadgroup (Metal packs thread_index_in_threadgroup
0-31 into simdgroup 0, 32-63 into simdgroup 1), so sharpemu_lane becomes
thread_index_in_threadgroup & 63 and cross-lane ops that span the full
wave rendezvous the two halves through threadgroup scratch:

- ballot into EXEC/VCC/SGPR pairs: each half's simd_ballot is written to
  its scratch slot, a threadgroup_barrier syncs, and all lanes recombine
  the 64-bit mask into the low/high register pair (centralized in
  EmitBallotStore, which the wave32 path shares).
- read-first-lane: broadcasts the lowest active lane's value across both
  halves through a scratch slot (EmitWave64ReadFirstLane).
- mbcnt lo/hi: 64-lane thread-mask math (no cross-lane op, just correct
  per-lane masks; lanes >= 32 would overflow a 32-bit shift, so split).

The barriers are safe because the guest's scalar PC keeps all 64 lanes
lockstep through the dispatcher. Scope matches the SPIR-V reference: the
scratch is indexed by half, so correct for a one-wave (64-thread)
workgroup, and readlane across halves stays a 32-wide shuffle. Wave-
agnostic wave64 kernels still translate per-thread unchanged.

Verified on the real GPU (MetalRuntimeTests): the emitted wave64 MSL
compiles, and a 64-lane dispatch runs through the bridge barriers
without deadlocking, returning the broadcast value. All 27 tests pass;
Dreaming Sarah (60/60) and void Terrarium (in-game, 58 flips) show the
shared wave32 ballot path is unaffected.

* [Gpu] Metal samplers: bind through an argument buffer, lifting the 16-slot cap

Metal exposes only 16 direct [[sampler(N)]] slots per stage, but real
shaders sample more (void Terrarium's scene shader: 17 images) and were
dropped at translation. Route samplers through a per-stage argument
buffer instead: the MSL declares a Gen5Samplers struct (one sampler per
sampled image, [[id(N)]]) taken as constant& at a buffer slot past the
stage's globals/uniforms/scalar-state, and the runtime writes each
sampler's Tier 2 gpuResourceID into an arena slice bound there. Textures
stay on direct [[texture(N)]] slots (31 is enough). One sampler per
image keeps them distinct, matching the SPIR-V/Vulkan path — no dedup,
so no wrong-sampler artifacts.

Verified argument buffers lift the limit on Apple Silicon (20-sampler
pipeline probe). Dreaming Sarah renders correctly at 60/60 with Metal
API validation clean; the void Terrarium scene shader that exceeded the
limit now compiles and runs (draws 74 -> 102/frame); all 27 shader
compiler tests pass, goldens unchanged (fixtures sample nothing).

* [Gpu] Metal: Shared storage for CPU-populated, GPU-sampled textures

The MTLTextureDescriptor default is Managed, which on unified memory needs
an explicit host->device sync we never issue after replaceRegion, so the
GPU can sample stale texels. These textures are CPU-uploaded and GPU-read,
so Shared (coherent, no sync on Apple Silicon) is the correct mode.

* [HLE] Add missing AGC/AudioOut/Pad exports blocking Unity+FMOD titles

Four exports were unresolved and hard-stalled GPU/audio/input init in
Unity titles (Lunar Lander Beyond froze there before opening VideoOut):

- sceAgcDriverSetTFRing / sceAgcDriverSetHsOffchipParam: tessellation-ring
  and hull-shader off-chip config. We translate shaders directly, so these
  only need to report success for init to proceed.
- sceAudioOutGetPortState: report a connected primary output at full volume.
- scePadDeviceClassGetExtendedInformation: report a standard pad (no special
  peripheral) so device-class probes resolve.

Generic HLE, backend-agnostic (helps the Vulkan path equally).

* [VideoOut] RegisterBuffers2: mask the 32-bit category, accept COMPRESSED

sceVideoOutRegisterBuffers2's category is a 32-bit SceVideoOutBufferCategory
passed on the stack, but we read the full 64-bit slot — whose upper word
carries stale GNM magic (0xC0DEC0DE...) the caller never cleared. The old
check then rejected every call as INVALID_VALUE, so buffer registration
failed and no frame ever presented. Mask to 32 bits and accept both
UNCOMPRESSED (0) and COMPRESSED (1); we present either identically.

Fixes Lunar Lander Beyond reaching its window (now presents 3840x2160).

* [HLE] Stub sceAudioPropagation (3D-audio) so Astro Bot boots past its assert

Astro Bot hard-crashed right after the splash: it calls
sceAudioPropagationSystemQueryMemory during audio init, and because the
whole libSceAudioPropagation module was unimplemented the call failed, so
the game asserted (AudioPropagationContext.cpp:43) and executed int 0x41 to
abort — an unrecoverable trap that kills the process.

We don't model acoustic propagation (geometry-driven reverb/occlusion is a
quality feature, not a correctness gate). The API is placement-style, so
QueryMemory reports a buffer size and the rest succeed as no-ops: the system
lives in the caller's own buffer. All 39 entry points stubbed; the game now
boots past the assert to the presenter. Backend-agnostic HLE.

* [Kernel] pthread_cond_wait: don't spuriously EPERM an untracked mutex

pthread_cond_wait/timedwait required our host-side mutex tracking to show
the calling thread as the owner, else it returned EPERM. But libkernel's
uncontended mutex fast-path locks the mutex word in guest memory directly,
without an HLE call, so we often never observe the lock and see owner==0.

Real pthread_cond_wait requires the caller to hold the mutex but does not
verify it for normal mutexes, so EPERM here is doubly wrong: it spins the
guest (Hades hammered this millions of times/sec) and, worse, skips the
unlock — leaving the mutex held and wedging every thread that later blocks
on pthread_mutex_lock. When the mutex reads as untracked (owner==0), adopt
ownership so the unlock/wait/re-lock cycle is balanced and actually releases
it. Genuine ownership violations (owned by another thread) still error.

Eliminates the EPERM storm and converts the resulting livelock into correct
blocking; no effect on games that lock through the HLE (owner already set).

* [Core] SSE4a EXTRQ patch: read the xmm register from ModRM, not xmm2

The loader rewrites Sony's AMD-only SSE4a EXTRQ+blend idiom into SSE4.1 at
boot, because Rosetta 2 and Intel hosts raise #UD -> SIGILL on EXTRQ. The
matcher hard-coded the source register to xmm2 (ModRM 0xC2), but the compiler
allocates it freely: Dead Cells (PPSA15552) emits the identical idiom against
xmm1, so it slipped through unpatched and the game died with SIGILL right
after the first frame.

Read the register from the ModRM r/m field instead, covering xmm0-xmm7, and
require it to be consistent across the EXTRQ and the blend. The pure
match/encode logic is extracted into Sse4aExtrqBlendPatch, isolated from the
native page-patching, and unit-tested for every register plus the round trip
and rejection cases; DirectExecutionBackend just applies it.

Dead Cells now patches its xmm1 idioms and boots past the first frame.

* [Ngs2] Implement non-allocator sceNgs2SystemCreate / sceNgs2RackCreate

Dead Cells uses the non-allocator NGS2 create entry points, which were
unimplemented. sceNgs2SystemCreate came back as an unresolved import, so the
game got a garbage system handle; every downstream sceNgs2RackCreate /
sceNgs2RackGetVoiceHandle then failed, the voice handle stayed null, and once
gameplay started the audio path polled sceNgs2VoiceGetState/VoiceControl on
the null voice forever — freezing the game in-level at FLIP 0.

The non-allocator forms differ only in a caller buffer (rsi/rcx) vs an
allocator callback; the system/option and out-handle arguments sit at the
same positions, so they alias the existing WithAllocator implementations.
Resolves the NGS2 InvalidVoiceHandle storm (591+/run -> 0).

* [SaveData] Real save subsystem: ~/SharpEmu/Saves/<titleId>, events, full CRUD

Rework the SaveData HLE from a partial stub into a working subsystem:

- Storage moves to ~/SharpEmu/Saves/<titleId>/<dirName>/ (was next to the
  exe under user/savedata/<userId>/<titleId>), overridable via
  SHARPEMU_SAVEDATA_DIR. Metadata (title/subtitle/detail/userParam) and icon
  live in <slot>/sce_sys/. Pure path + param.json logic is isolated in a new
  SaveDataStorage type and unit-tested.
- Async event model: sceSaveDataGetEventResult now resolves (was an
  unresolved import a save worker polled forever), returning queued completion
  events or a clean 'no event' status; SyncSaveDataMemory posts a
  SAVE_DATA_MEMORY_SYNC_END event. Plus GetEventInfo/SetEventInfo/register
  callbacks.
- New exports: Mount/Mount2/Mount5/Umount, Delete/Delete5, GetParam/SetParam,
  SaveIcon/SaveIconByPath/LoadIcon, GetAllSize/GetProgress/GetMountInfo/
  IsMounted/GetSaveDataCount/GetMountedSaveDataCount/Abort, Initialize/
  Initialize2/Terminate, SaveDataMemory v1 aliases.
- Mounts are tracked so Umount2 really unregisters the /savedata0 mapping
  (new KernelMemoryCompatExports.UnregisterGuestPathMount) and params/icons
  resolve against the live mount; DirNameSearch surfaces param.json titles.

15 new unit tests (storage layout/sanitize/metadata + mount/event/param/delete
exports); full suite 277 passing.

* [Gpu] Metal: Cmd+F1 toggles Apple's Metal Performance HUD

Plain F1 keeps the built-in CPU-rasterized perf overlay; Cmd+F1 now toggles
the system Metal Performance HUD on the CAMetalLayer, Metal backend only.

Command-modified keys never reach keyDown: (AppKit routes them through the
key-equivalent chain), so the input view gains a performKeyEquivalent:
override that claims Cmd+F1 (also silencing the system beep) and leaves
everything else to the responder chain.

Configured per Apple's 'Customizing Metal Performance HUD':
developerHUDProperties with mode=default + logging=default, plus
MTL_HUD_LOG_SHADER_ENABLED=1 passed directly in the dictionary — HUD,
per-frame statistics logging, and shader-compile logging all enabled from
one property set; mode=disabled hides it again. Guarded by a
respondsToSelector: check for older macOS.

* [Gpu] Metal: also catch Cmd+F1 in keyDown: for the HUD toggle

Function keys reach keyDown: even with Command held (AppKit only reroutes
some chords through performKeyEquivalent:), so the HUD toggle was never
firing there. Handle Cmd+F1 in both the keyDown: and performKeyEquivalent:
paths, and keep it out of MetalHostInput so it can't also flip the plain-F1
perf overlay.

* [Audio] Diagnostics: NGS2 voice-param dump + AudioOut peak-amplitude trace

Two gated traces (idiomatic SHARPEMU_LOG_* style) that pinpoint where audio
dies for NGS2-based games:

- SHARPEMU_LOG_NGS2 now walks the sceNgs2VoiceControl param list and logs each
  {size,id} block header + payload bytes, confirming the real layout
  (header = u32 size, u32 id; waveform-block param id=0x10000001 carries the
  guest PCM pointer at +8; rate param id=0x10000005 carries the resample ratio).
- SHARPEMU_LOG_AUDIO_OUT logs sceAudioOutOutput call count and the peak
  amplitude of each submitted buffer.

Finding on void Terrarium: sceAudioOutOutput is called thousands of times on
both 8ch/float32 ports, but every buffer has peak=0.0 — the guest submits pure
silence. The host path (AudioOut -> PCM convert -> CoreAudio) is proven correct;
the silence originates in Ngs2SystemRender, which zeroes the output buffer
instead of mixing voices. Restoring audio for NGS2 games requires a real NGS2
software mixer (next).

* [Audio] NGS2 software mixer: decode + mix PS-ADPCM voices

NGS2-based games were silent because sceNgs2SystemRender only zeroed the
output buffer. This adds a real software mixer:

- Ngs2VagDecoder: clean-room PS-ADPCM ("VAGp") decoder producing mono PCM16
  with loop points resolved from the exact per-frame flag values (3=loop
  start, 6=loop end, 1/7=one-shot end).
- Voice control now parses the SceNgs2VoiceParamHead command list, decodes the
  waveform-blocks param's VAGp container once, and arms the voice.
- sceNgs2SystemRender mixes every armed voice belonging to the system into the
  leading grain of the render buffer as interleaved float32 (nearest-sample
  resample from the source rate to 48 kHz, additive into the front L/R pair),
  which is exactly what games copy to sceAudioOutOutput.

Verified on void Terrarium: previously peak=0.0 silence at AudioOut, now real
audible SFX/music. Voices are still armed on waveform assignment rather than an
explicit kick, so pooled/duplicate voices can overlap — trigger-state handling
is a follow-up.

* [Gpu] AGC: latch GPU-wait satisfaction to the produced value

Fixes a lost-wakeup race that stalled games at a black/splash screen. When a
RELEASE_MEM packet writes a completion label, the guest frequently resets that
label to 0 immediately to reuse it next frame. Our wake path
(GpuWaitRegistry.CollectSatisfied) re-reads *current* guest memory, so if the
reset lands before the wake pass runs, the transient satisfied window is missed
and the suspended DCB waits forever — even though the producing write executed
(traced as wrote=True) and its producer is marked completed.

RELEASE_MEM producers now call GpuWaitRegistry.LatchSatisfiedByValue with the
value they actually wrote, recording satisfaction at the moment of the write for
any waiter that value satisfies. CollectSatisfied honors the latch regardless of
the current (possibly-reset) memory value. This is fail-closed: a waiter only
latches when a real producer wrote a genuinely satisfying value.

Verified: Astro Bot's DEADBEEF sentinel wait (dcb.graphics waiting on a
release_mem label) that was permanently stuck is now resolved; void Terrarium is
unregressed (runs, audio intact, no producerless stalls). Astro still has
separate unresolved blockers (producer-behind-its-own-wait cascades and
producer=none-observed labels) tracked for follow-up; WRITE_DATA/DMA_DATA
producers could latch too but are left out until there is evidence they race.

* [Gpu] AGC: retry indirect dispatches whose GPU-computed dims aren't ready

GPU-driven games (Astro Bot) build their frame on the GPU: a compute dispatch
writes the thread-group dimensions for the next DISPATCH_INDIRECT into a guest
buffer. Our AGC parser reads those dimensions on the CPU at parse time, which
runs before the producing dispatch has executed on the render thread — so it
read 0/0/0 and dropped the work (agc.dispatch_reject zero-dimension), leaving
the scene unrendered (black) and cascading into stuck cross-queue fence waits.

Instead of dropping a zero-dimension INDIRECT dispatch, suspend the DCB on its
dimensions buffer (reusing the WAIT_REG_MEM suspend/resume + GpuWaitRegistry
machinery) until the producer writes non-zero dims, then re-parse and dispatch.
A bounded per-wait deadline (150 ms) resumes-and-drops a genuinely empty
indirect dispatch so it can never stall the queue, making the change
non-regressive: worst case matches the old drop behavior after a short wait.
Direct dispatches (dims inline) are unaffected.

Result: Astro Bot goes from a permanent black screen to actually rendering
(the presenter reports "Metal VideoOut presenting 3840x2160"). void Terrarium —
which issues no indirect dispatches — is unregressed (runs, audio intact, zero
rejects). Astro then hits a separate, newly-reached downstream crash (guest
TBB worker thread_set_state failure) tracked for follow-up.

* [ShaderCompiler] Metal: keep compute shaders within read_write and LDS limits

Two Metal limits made real Astro Bot compute shaders fail to compile/create,
which dropped their dispatches and cascaded into stuck GPU waits (splash hang):

- Textures with access::read_write are capped at 8 per function, but every
  storage image was declared read_write. Track each binding's actual access
  during body emission (ImageLoad->read, ImageStore->write, ImageAtomic and a
  load+store sharing one binding->read_write) and emit the minimal qualifier,
  so read-only/write-only storage images no longer count against the cap.

- Threadgroup memory is capped at 32 KB. A shader requesting the full 32 KB of
  LDS plus the separate 3-dword wave64 bridge overflowed by 12 bytes. Alias the
  bridge into the top of the LDS allocation when both are used, mirroring the
  SPIR-V translator's _waveScratchInLds path, keeping the total at 32 KB.

Verified on Astro Bot: "read_write access exceeds maximum (8)" and "Threadgroup
memory size (32780) exceeds maximum (32768)" are both gone; the 27 MSL golden
tests still pass (no golden used a storage image or LDS+wave64 shader).

* [Gpu] AGC: break cross-queue GPU wait deadlocks with a produced-value fallback

Real GPU-driven titles (Astro Bot) drive graphics and compute queues with
mutually dependent WAIT_REG_MEM fences: graphics waits on a compute EOP label,
compute waits on a graphics label. On hardware the two queues run concurrently
so the cycle resolves, but our submission parser is serial, so a label that gets
written -> reset for reuse -> re-waited across queues can wedge forever. The
latch fix helped the write-then-consume race but the cycle re-formed each frame
(graphics stuck at 3 flips, compute queues permanently suspended).

Producers now record the last value they wrote to each label
(GpuWaitRegistry.RecordProduced). A new deadlock breaker
(CollectDeadlockBroken, run from DrainResumableDcbs) releases any waiter stuck
past a 500 ms deadline whose condition is satisfied by that recorded value —
i.e. a real producer signalled the label at least once, guest memory has just
since been reset. It never fabricates a value, and the long deadline means
legitimate fences (which complete within a frame) never trip it.

Verified: Astro Bot goes from 3 flips (wedged on splash) to 25, loads its
splash level ("LevelDocument Loaded: ps_logo") and produces 2432x1368 frame
content. void Terrarium is untouched — 0 deadlock-break events, 1020 flips,
audio intact (its waits resolve far under the deadline). Tunable via
SHARPEMU_GPU_DEADLOCK_BREAK_MS.

* [Cpu] SSE4a EXTRQ patch: cover any blend destination register, not just xmm0

The EXTRQ+VPBLENDD idiom rewrite only matched when the blend destination was
xmm0 (VEX.vvvv byte 0x79). Sony's toolchain allocates that register freely: a
Dead Cells build emits `EXTRQ xmm4,0x28,0x00 ; VPBLENDD xmm3,xmm3,xmm4,2`
(VEX byte 0x61, dest xmm3). That instance stayed unpatched, so the AMD-only
EXTRQ reached Rosetta 2 and raised #UD -> SIGILL (0xC000001D) the moment the
game entered gameplay (loading level PrisonStart).

Read the destination register from VPBLENDD's VEX.vvvv / ModRM.reg as well as
the source from the ModRM r/m field, and emit PINSRD into that destination. Both
are still constrained to xmm0-xmm7 by the fixed VEX prefix. Match/encode stay in
the unit-tested helper.

Verified: Dead Cells now patches 14 EXTRQ blends (previously 0 on this build),
no SIGILL, and reaches PrisonStart. 15 patch unit tests pass, including the exact
xmm3/xmm4 bytes that faulted.

* [HLE] Implement Dead Cells' remaining unresolved imports

Three imports Dead Cells calls during boot/level-load were unresolved, so they
returned no defined value:
- scePadGetHandle (libScePad): returns the primary pad's handle (polled every
  frame for input); same validation as scePadOpen.
- sceNpEntitlementAccessGetAddcontEntitlementInfo (libSceNpEntitlementAccess):
  singular add-on-content lookup; we own no DLC, so zero the info out and return
  OK, matching the existing list variant.
- sceNpUniversalDataSystemEventPropertyArraySetString: telemetry setter, dropped.

Dead Cells now boots with zero unresolved imports. (It still stalls later at
PrisonStart level-load — a separate GPU/threading issue, not an import gap.)

* [Kernel] Fix pthread mutex deadlock: trylock semantics + stale-waiter clog

Hades hard-froze during boot on a "free but reserved" mutex: owner==0 yet
every acquisition failed forever. Two independent defects in the pthread
mutex compat layer combined to wedge it, both traced from real runs.

1. trylock incorrectly required an empty wait queue. POSIX
   pthread_mutex_trylock succeeds whenever the mutex is not currently held
   and owes no fairness to queued waiters; gating it on Waiters.Count==0 made
   a spin-on-trylock loop (which the game runs) spin forever against a single
   undrainable waiter even though owner==0. trylock now acquires on owner==0;
   the blocking lock still honours FIFO so genuine blocked waiters are not
   starved by a barging locker.

2. cond_timedwait timeouts leaked mutex re-acquire waiters. A cond wait's
   timeout enqueues a re-acquire waiter whose wake hand-off can be lost,
   orphaning it in the mutex queue. Multiple orphans from one thread piled at
   the FIFO head; the unlock hand-off then woke a dead wake-key and the mutex
   never drained. A thread can hold at most one pending acquisition on a
   mutex, so EnqueueMutexWaiterLocked now prunes any prior waiter for the same
   thread before enqueueing — collapsing the leaked pile.

Verified: Hades advances from a hard freeze at ~4.9M HLE calls (main and a
worker both blocked on the same free mutex) to 24.9M calls with no stall,
reaching the save-data/user-service boot stage. void tRrLM behaves
identically with and without the change (no regression); all 268 Libs tests
pass.
2026-07-18 20:32:00 +03:00
Spooks 13269797bf Add live debugger frontend and mutex stall recovery (#383) 2026-07-17 22:41:07 -06:00
Peter Bonanni 0b1dea43e8 [CPU] Preserve blocked leaf import waiters (#350) 2026-07-18 02:59:13 +03:00
Berk fbafd3f429 [GUI] Add native Vulkan host surface support and more (#337)
* [GUI] Add native Vulkan host surface support and more

* reuse

* [GUI] set width session menu
2026-07-17 18:57:10 +03:00
Mees van den Kieboom aa25f6e978 [Loader] Restore PS5 SELF header support (#342)
Accept both PS4 and PS5 SELF signatures without treating version, key type, or flags as layout markers. Add synthetic coverage for valid header variants and malformed structural fields.

Signed-off-by: missatjuhvdk1 <177474143+missatjuhvdk1@users.noreply.github.com>
Co-authored-by: missatjuhvdk1 <177474143+missatjuhvdk1@users.noreply.github.com>
2026-07-17 18:55:16 +03:00
kuba 28485b60e2 CPU: avoid continuation emitter closure allocations (#306) 2026-07-17 04:17:58 +03:00
Tell-Shanks 0755ca15f7 core: report what occupies a fixed-address allocation on failure (#278)
* Implement DescribeAddressForDiagnostics method

Added a method to describe the state of a memory address for diagnostics.

* core: enrich allocation failure exception with host region diagnostic
2026-07-17 02:58:31 +03:00
kuba 33f96252da [CPU] Reject context transfers to unmapped guest addresses (#273) 2026-07-16 19:57:34 +03:00
MikeyLITE69 52d2874fa8 cpu: emulate BMI1/BMI2/ABM instructions in software when the host lacks them (#249)
## What

The native backend runs guest code directly on the host CPU. When the host doesn't
implement a BMI1/BMI2/ABM instruction that the PS5's Zen 2 cores do, it raises #UD
(STATUS_ILLEGAL_INSTRUCTION). Today the vectored handler just logs the faulting bytes
and gives up, so the title dies.

This adds a software fallback. On an illegal-instruction fault we decode the opcode
with Iced (the decoder already used elsewhere in the backend), evaluate it against the
trapped register/memory state, write the result and flags back into the CONTEXT record,
step RIP past the instruction, and resume.

Instructions covered (32- and 64-bit): ANDN, BLSI, BLSMSK, BLSR, BEXTR, BZHI, TZCNT,
LZCNT, RORX, SARX, SHLX, SHRX, PDEP, PEXT.

## Why

Users on CPUs without these extensions currently can't get past code that uses them.
This is a generic fix (no game-specific hacks) that improves compatibility on older
hosts. MULX is intentionally left out for now — its dest_hi/dest_lo operand ordering
is easy to get subtly wrong, so I'd rather add it separately with its own tests.

## How it's structured

- `BmiInstructionEmulator` holds the pure bit/flag semantics with no dependency on the
  unsafe CONTEXT plumbing, so it can be unit-tested directly.
- `DirectExecutionBackend.IllegalInstruction.cs` is the thin unsafe adapter (decode →
  read operands → emulate → write back → advance RIP). Anything it doesn't fully model
  returns false and falls through to the existing diagnostics unchanged, so it can never
  mis-handle an opcode it doesn't recognize.
- One hook in `DirectExecutionBackend.Exceptions.cs`, next to the other TryRecover* calls.
- Emits a single one-time "emulating in software" log line, not per-instruction spam.

## How I verified

- Added xUnit tests covering every instruction in both widths plus the CF/ZF/SF/OF
  edge cases (src == 0, shift-count masking, index beyond operand width, etc.).
- Cross-checked all the expected values against an independent reference implementation
  written from the Intel/AMD definitions; results match.
- `dotnet build` + `dotnet test` pass locally.

## Notes

New files follow .editorconfig (4-space, SPDX headers, REUSE-compliant).
2026-07-16 16:38:27 +03:00
Gutemberg Ribeiro c5a82c1065 [Perf] Behavior-preserving hot-path allocation/LINQ wins (#264)
* [Perf] Gate event-flag tracing so it allocates nothing when disabled

TraceEventFlag built its interpolated argument (and, on the wait path,
FormatFrameChain + a new StringBuilder(256) in FormatGuestWaitObject plus
~12 guest-memory reads) on every call, then checked the env var inside the
method — so every sceKernelSetEventFlag/Clear/Poll/Wait paid a string
allocation and an Environment.GetEnvironmentVariable P/Invoke even with
tracing off. Cache the flag once in a static readonly bool and gate every
call site, matching the semaphore/event-queue pattern. Behavior is
unchanged when SHARPEMU_LOG_EVENT_FLAG=1.

* [Perf] Hoist IsNoBlockLeaf classification to import-stub setup

The leaf-dispatch path called IsNoBlockLeafImport(nid) — a ~30-literal
string pattern match — on every leaf import. IsLeaf/NidHash were already
precomputed on ImportStubEntry at stub setup; add IsNoBlockLeaf alongside
them and read the field in the hot path. Behavior unchanged.

* [Perf] Cache trace env-var flags read on hot paths

Two SHARPEMU_LOG_* env vars were read via Environment.GetEnvironmentVariable
(a P/Invoke + transient string) on hot paths: SHARPEMU_LOG_FIBER on every
fiber context transfer, and SHARPEMU_LOG_DIRECT_MEMORY on every direct-memory
op (~8 sites). Cache both once — _logFiber alongside the other backend _log*
flags, and _traceDirectMemory behind the existing ShouldTraceDirectMemory
helper. Behavior unchanged.

* [Perf] Avoid per-iteration thread snapshot in the idle pump loop

PumpUntilGuestThreadsIdle allocated a full GuestThreadState[] snapshot
(via LINQ Values.ToArray()) on every spin just to tally three run-state
booleans. Tally them under the lock with an allocation-free helper, and
only materialize the snapshot inside the gated (default-off) diagnostic
dump. SnapshotGuestThreads now uses Values.CopyTo instead of LINQ.
Behavior unchanged.

* [Perf] De-LINQ GetPixelColorExportMask on the per-draw path

GetPixelColorExportMask ran a Select/OfType/Where/Aggregate chain over all
shader instructions, allocating iterators + closures, and is called per
render target (twice per draw via CreateRenderState and HasPixelColorExport).
Replace with a manual scan producing the identical mask — no allocation, and
it removes an authored-LINQ use the repo bans.

* [Perf] De-LINQ per-draw render-target selection in AgcExports

The bound-render-target selection used Where(...).OrderBy(...).ToArray()
(plus a second Where/ToArray fallback) on every translated draw, allocating
LINQ iterators + closures. Replace with an explicit filter into a pre-sized
list + List.Sort by slot; slots are distinct so this matches the stable
OrderBy. Same result, no per-draw LINQ allocations.

* [Perf] De-LINQ per-draw render-target validation in the Vulkan presenter

SubmitOffscreenTranslatedDraw validated its render targets with
targets.Any(...) twice plus targets.Select(a=>a.Address).Distinct().Count()
(a per-draw HashSet), and broadcast a single blend with
Enumerable.Repeat(...).ToArray(). Targets are <= 8, so replace with manual
scans (invalid-target check; combined dimension-mismatch + pairwise
aliasing check) and Array.Fill. Same results, no per-draw LINQ allocations.

* [Perf] Binary-search VirtualQuery region lookup (SortedList)

_mappedRegions was an unordered Dictionary, so TryFindVirtualQueryRegionLocked
scanned every region for containment/next — O(n) per sceKernelVirtualQuery and
O(n^2) when an allocator walks the address space with the findNext flag. Store
regions in a SortedList keyed by base address (every write already uses the
region's own address as the key) and find the containing/next region with a
binary search over the sorted keys. Also drops a now-redundant Values.OrderBy.

Non-overlapping regions assumed (mmap semantics), so only the floor region can
contain the query. Behavior-preserving; worth spot-checking VirtualQuery-heavy
titles.

* [Perf] Remove stray CLI packages.lock.json committed by mistake

git add -A in an earlier commit swept in a regenerated
src/SharpEmu.CLI/packages.lock.json. main tracks no lock files (central
package management, no RestorePackagesWithLockFile), and REUSE.toml no
longer covers packages.lock.json, so the committed file failed the REUSE
Compliance check. Remove it.
2026-07-16 16:37:28 +03:00
Spooks b4b95014f1 Fix/deadcells crash (#262)
* Boot compatibility fixes for UE titles, GUI toggles, and DeS render/boot work

Checkpoint of the Monster Truck Championship and Demon's Souls boot work.
Each piece is independently useful and verified against the titles.

- playgo: scePlayGoGetLocus now returns BAD_CHUNK_ID for chunk ids outside
  the known set, matching real firmware. Titles enumerate chunk ids until
  that error; answering OK for every id made the scan wrap the ushort range
  and spin forever. Missing-sidecar and no-app0 fallbacks report a
  fully-installed single chunk 0 so scePlayGoOpen keeps succeeding.
- kernel: restore the SHARPEMU_WRITABLE_APP0 opt-in. Unpackaged UE dumps
  write their Saved tree under /app0 during PS5 component init and treat
  the denial as a fatal boot error.
- pad: accept handle 0 as the primary pad across all pad calls. Real
  firmware hands out small non-negative handles and some titles read state
  with handle 0.
- bthid: env-gated experiment hooks for the Thrustmaster wheel middleware
  investigation (fail-only-RegisterCallback modes and a synthetic
  enumeration callback with a zeroed event struct). All default off.
- gui: add SHARPEMU_LOG_IO and SHARPEMU_WRITABLE_APP0 toggles to the
  Environment tab.
- videoout: per-swapchain-image render-finished semaphores (the shared
  semaphore raced the swapchain); whole-mip-chain layout init for offscreen
  guest images (sampled binds read mips stuck in Undefined); GPU-resident
  texture availability now canonicalizes through the texture format table
  and accepts compatibility-class aliases, cutting per-frame CPU texture
  re-reads (143 GB -> 55 GB per 300 s in Demon's Souls, 0.2 -> 0.5 fps).
- hle: add sceSystemServiceGetNoticeScreenSkipFlag,
  sceSystemServiceGetMainAppTitleId (title id published from the runtime),
  and sceNpWebApi2CreateUserContext (refuses so the online layer backs off).
- rtc: SHARPEMU_RTC_PROBE_RANGE diagnostic dumps the code around a busy-wait
  caller of sceRtcGetCurrentTick once; costs nothing when unset.

* [ShaderCompiler] Fix VReadlaneB32 scalar destination field

The scalar destination lives in the low vdst byte (bits 0-7); it was read
from bits 8-14, the VOP3B carry-out field readlane does not have, sending
every readlane result to s0. Verified against raw gfx10 encodings and
LLVM's assembler tests (v_readlane_b32 s5, v1, s2 -> low byte 0x05).

* [VideoOut] Survive device loss and flip-order asserts without dying

Two ways a frame could take down the whole presenter:

- Device loss between any two Vulkan calls in a frame unwound the window
  thread, and the Dispose-time fence check then threw again, masking the
  original error. Catch the loss at the frame boundary, retire
  presentations and guest submissions whose fences can never signal, and
  keep the window loop pumping so the game (audio, logic) carries on.
- The ordered-flip capture invariant is violated ~100 times per run by
  Demon's Souls (PPSA01342); on debug builds the Debug.Assert fail-fasts
  the process with nothing in the log. Downgrade it to a once-per-version
  warning until the capture/wait ordering is understood.

* Fix Dead Cells shader cache regression

---------

Co-authored-by: StealUrKill <35749471+StealUrKill@users.noreply.github.com>
2026-07-16 13:57:16 +03:00
Spooks 9bacb883f1 Fix Linux aligned mapping retention (#247)
* Fix Linux aligned mapping retention

* Cover Linux aligned mapping retention
2026-07-15 19:34:35 -06:00
Miguel Cruz 864cbb0fa0 [AGC/Vulkan] Extend PS5 runtime and rendering compatibility (#216)
* [Core] Add POSIX native execution and PS5 SELF support

Extend the native backend, guest TLS, fixed-address memory, and loader paths needed by PS5 titles on Windows, Linux, and macOS. Keep workstation GC so high-core-count hosts do not reserve over fixed guest image bases.

* [HLE] Expand PS5 service and media compatibility

Add the kernel, threading, save-data, networking, audio, video-codec, font, dialog, and service exports required by newer PS5 software. Preserve every SysAbi NID currently registered by main while adding the compatibility surface used by ASTRO BOT.

* [AGC/Vulkan] Extend Gen5 shader and presentation support

Expand PM4 handling, Gen5 shader translation, MRT and packed export support, guest image tracking, depth initialization, texture aliasing, and Vulkan presentation. Add the performance overlay and address-filtered diagnostics used to validate ASTRO BOT with original shaders.

* [Core] Align static TLS reservation across hosts

* [Pad] Align primary user ID with UserService

* [Gpu] Preserve runtime scalar buffers across renderer seam

* [AGC] Restore omitted command helper exports

* [Vulkan] Reuse primary views for promoted MRT targets

* [Vulkan] Preserve scratch storage bindings in compute dispatches
2026-07-16 02:02:34 +03:00
Gutemberg Ribeiro 320dbcacba [SourceGenerators] Compile-time SysAbi export registry, analyzers, and build-generated aerolib.bin (#204)
* [SourceGenerators] Add the SysAbi export generator and analyzers (phase 0)

New SharpEmu.SourceGenerators Roslyn component, complete and tested but
consumed by nothing yet — the emulator projects adopt it in the
following commits.

Ps5Nid ports the PS NID derivation (base64 of the byte-reversed first
eight SHA1 bytes of name + fixed suffix) from
scripts/generate_aerolib_binary.py to C#, so what has always been a
manual, out-of-band computation becomes a compile-time capability.

SysAbiExportGenerator emits a per-assembly SysAbiExportRegistry whose
CreateExports(Generation) reproduces ModuleManager's reflection scan
exactly — same generation inheritance and filtering, same method-name
fallback, same libKernel default — with attribute-omitted NIDs derived
algorithmically (equivalent to the runtime catalog lookup, which is
built from the same computation). Parameterless handlers are adapted to
the SysAbiFunction shape; invalid declarations are skipped here because
the analyzer rejects them as build errors, so nothing drops silently.

SysAbiExportAnalyzer turns the runtime failure modes into diagnostics:
SHEM001 duplicate NID (across declared and derived forms), SHEM002
malformed NID, SHEM003 uncallable handler signature, SHEM004 NID
contradicting its export name (the class of drift previously fixed by
hand), SHEM005 unresolvable export, SHEM006 export name unknown to
ps5_names.txt when the catalog is wired as an AdditionalFile, SHEM007
handler not reachable by generated code.

The self-contained test suite drives both in-process against the real
SharpEmu.HLE metadata: known catalog NID pairs pin the algorithm, the
generated registry must itself compile, and each diagnostic has a
triggering fixture. Fittingly, the NID pinning test caught a wrong
pair in its own first draft — the exact mistake SHEM004 exists to stop.

* [SourceGenerators] Adopt the generated export registry in the emulator (phase 1)

SharpEmu.Libs consumes the generator and analyzers, with
scripts/ps5_names.txt wired as the AdditionalFile catalog. The runtime
now registers exports from the compile-time SysAbiExportRegistry
instead of the boot-time reflection scan; RegisterFromAssembly is
retained solely as the arbiter for a parity test that pins the two
tables identical — same NIDs, names, libraries, targets, and handler
methods — across Gen4, Gen5, and combined registration.

First contact between the analyzer and all 715 existing exports
surfaced real drift the old offline checker structurally missed
(scripts/check_sysabi_aerolib.py skipped any NID absent from
aerolib.bin): three exports whose friendly names collide with real
catalog symbols of different NIDs, now suppressed at-site with reasons
pending AGC API confirmation, alongside the established synthetic
Unknown* labels for uncatalogued NIDs, which prompted a rule
refinement — SHEM004 only hard-errors when the export name is a real
catalog symbol, since synthetic labels cannot be validated by hashing
and the NID is authoritative for them. The two allowlisted mismatches
in the python checker no longer trigger anything, and the checker is
deleted: the analyzer subsumes it with the semantic model instead of
regex, and validates every declared pair rather than only
catalog-known NIDs.

* [SourceGenerators] Generate aerolib.bin at build time from ps5_names.txt

The runtime NID -> name catalog is derived data and no longer lives in
the repository: a Framework-only MSBuild task (GenerateAerolibBinaryTask,
sharing the same Ps5Nid implementation the analyzers use) builds it
into the intermediate directory from scripts/ps5_names.txt — now the
single source of truth — and SharpEmu.HLE embeds it from there. The
output is byte-identical to the previously committed binary, verified
with cmp against git history; a new test pins that the embedded catalog
loads and resolves a known symbol both directions.

scripts/generate_aerolib_binary.py is deleted (its algorithm lives in
Ps5Nid, its invocation in the build); the REUSE annotation for the
binary goes with it. MSBuild's Inputs/Outputs check means the ~154k NID
hashes only recompute when the names file actually changes. The task
implements ITask against Microsoft.Build.Framework directly, keeping
the vulnerable-flagged Utilities.Core package out and the analyzer
project's file-IO ban suppressed only inside the task itself.

* [SourceGenerators] Emit typed-signature register thunks (phase 2)

[SysAbiExport] handlers can now be written with real signatures — a
CpuContext followed by up to six int/uint/long/ulong parameters — and
the generator emits the SysV unmarshalling thunk, mapping parameters
positionally to RDI/RSI/RDX/RCX/R8/R9 with the same unchecked-cast
idiom hand-written handlers use. SHEM003 accepts the new shape and
rejects register overflow and non-register-representable types. Both
shapes coexist, so migration is per-handler; sceKernelPollSema,
sceKernelSignalSema, and sceKernelCancelSema migrate as the
demonstration (the last showing raw ulong guest-address passthrough).

The reflection scan cannot represent typed handlers, so it retires
here: RegisterFromAssembly, its signature validation, and
ResolveExportInfo are deleted, and the parity test that pinned the
generated registry to the scan is replaced by content-invariant tests
(duplicate-free, full 715-export surface, catalog identity). Deleting
the scan surfaced a phase-1 latent regression — the pre-JIT warm sweep
enumerated only reflection-scanned assemblies, so the generated
registration path warmed nothing and re-exposed the guest-thread
fail-fast risk; the warm set is now derived from the registered
handler delegates themselves.

* [SourceGenerators] Marshal guest strings declaratively with [GuestCString] (phase 3)

A string parameter on a typed [SysAbiExport] handler, annotated
[GuestCString(maxLength)], now makes the generated thunk read the
null-terminated UTF-8 string from the argument register's guest
address before the handler runs, returning
ORBIS_GEN2_ERROR_MEMORY_FAULT to the guest when the read fails —
the exact prologue nearly every string-taking handler writes by hand.
The attribute lives in SharpEmu.HLE next to SysAbiExportAttribute;
SHEM008 rejects misuse (non-string parameter, non-positive MaxLength)
while a bare string parameter stays a SHEM003 signature error.

_open, open, and sceKernelOpen migrate as the demonstration; they were
chosen because their hand-written prologue faulted on a null pointer
the same way the thunk does (handlers that return INVALID_ARGUMENT for
null pointers, like sceKernelCreateSema, keep the raw shape so guest-
visible semantics stay untouched).

* [SourceGenerators] Apply review findings across the branch

Behavior: the open/_open/sceKernelOpen [GuestCString] demo migration is
reverted — the local compat reader falls back to host memory for paths
in loader-mapped regions that ctx.Memory cannot see, so the generated
thunk would have turned recoverable reads into MEMORY_FAULT. The
marshalling infrastructure stays, proven by generator/analyzer tests;
production migration waits for a handler whose semantics the thunk
reproduces exactly. A comment on the handler records why.

Build robustness: the aerolib target is skipped for design-time builds
(the IDE resolves project references without compiling them, so on a
fresh clone the task assembly does not exist yet), and the task/names
paths are centralized in properties. The generator now emits no
registry for export-free assemblies, so referencing the analyzer can
never mint a colliding SharpEmu.Generated type.

Cleanup and perf: the pragma-suppression sites left mis-indented by the
phase-1 relocation are reformatted and the restores moved after the
method body; the dead ExportsForTesting hook and its InternalsVisibleTo
are deleted; the aerolib task reuses one SHA1 instance across ~150k
names; the analyzer caches the parsed catalog per file snapshot instead
of re-parsing 150k lines every compilation start, shares the attribute
name constant with the generator, and computes the catalog-membership
check once.

* [CI] Run the test suites in the build workflow

The workflow compiled the test projects (they are in SharpEmu.slnx) but
never executed them. A solution-level dotnet test now runs between
build and publish, so any test failure fails the build — including the
AerolibCatalogTests/SysAbiRegistryTests that guard the build-generated
aerolib.bin and the generated export registry. Generation failures of
aerolib.bin itself already fail the build step: the MSBuild task logs
an error event and returns false, and a missing task assembly or
missing embedded output are hard MSBuild errors. The NuGet cache key
now also tracks the test projects' lock files.

* [SourceGenerators] Address review feedback

Multi-diagnostic analyzer tests no longer assume a stable diagnostic
order (analyzer execution is concurrent), and the aerolib task logs
the full exception instead of only its message so build failures keep
the type and stack trace.

* [SourceGenerators] Address second review round

Symbol-name comparisons in the shape rules and analyzer now pin an
explicit SymbolDisplayFormat.FullyQualifiedFormat instead of relying on
the display-format default, and the aerolib task fails loudly on a
symbol name that would overflow the format's ushort length prefix
instead of silently truncating it, with null-safe output-directory
handling made explicit.

* [SourceGenerators] Embed aerolib.bin via a target so design-time builds never reference it

The static EmbeddedResource item referenced the generated file even in
design-time builds, where the generation target is skipped — on a fresh
clone the IDE would try to embed a file that never existed. The item is
now created inside an EmbedAerolibBinary target gated on
DesignTimeBuild, separate from the generation target so an up-to-date
skip of GenerateAerolibBinary cannot drop the item with the rest of its
body, and hooked before AssignTargetPaths since dynamic resource items
added later miss the resource pipeline.

Verified fresh build, incremental rebuild (embedded catalog test both
times), and a simulated design-time compile with no artifacts present.

* [SourceGenerators] Regenerate test lock file after rebase onto main

Rebase fallout: main's package graph shifted under #200, so the
SourceGenerators.Tests lock file is re-evaluated to keep --locked-mode
restore green at the branch tip.

* [Build] Drop NuGet lock files; rely on central package management

Central package management was already in effect (ManagePackageVersionsCentrally
with all versions in Directory.Packages.props and no inline PackageReference
versions), so the per-project packages.lock.json files and the lock-mode
workflow only added maintenance overhead. This removes all eleven lock files,
drops RestorePackagesWithLockFile so restore no longer regenerates them, and
takes --locked-mode off the CI restore steps (re-keying the NuGet cache on the
central props files). Package versions remain centrally pinned in
Directory.Packages.props.
2026-07-16 00:00:32 +03:00
Gutemberg Ribeiro 30fdd8d6ed [Gpu] Backend-neutral shader compiler and guest-GPU renderer seam (#200)
* [ShaderCompiler] Extract the backend-neutral shader compiler project

Move the Gen5 (gfx10) microcode decoder, the scalar evaluator, the
shader IR, and the metadata reader out of SharpEmu.Libs/Agc into a new
SharpEmu.ShaderCompiler project — the half of shader compilation every
codegen backend (SPIR-V today; MSL and DXIL later) consumes. Types go
public: they are the contract now. Nothing in the project may depend on
a host graphics API; the SPIR-V-specific artifact types
(Gen5SpirvShader, Gen5SpirvStage) stay beside the emitter in Libs.

Three couplings surfaced by the move, each resolved at the right depth:
GuestDrawKind was defined inside VulkanVideoPresenter despite being a
guest-domain, decoder-produced concept — it moves to the shared project;
the evaluator's one HLE dependency (the tracked-libc-heap read
fallback) becomes an injectable hook that a Libs module initializer
installs before any caller can reach the evaluator; and the inline-
constant table is promoted to a shared Gen5InlineConstants so backends
cannot drift on constant semantics (the SPIR-V translator now delegates
to it).

The ShaderDump tool drops its reflection over the moved types in favor
of direct typed calls; only the SPIR-V emitter, still internal to Libs
until it moves to its own backend project, is reached via reflection.
Verified by a clean solution build, the existing test suite, and a full
ShaderDump conformance run.

* [ShaderCompiler] Move the SPIR-V emitter into SharpEmu.ShaderCompiler.Vulkan

Gen5SpirvTranslator (with its ALU partial), SpirvModuleBuilder,
SpirvFixedShaders, and the Gen5SpirvShader/Gen5SpirvStage artifact types
move whole from SharpEmu.Libs/Agc into the first per-backend codegen
project. Notably it needs no Vulkan bindings reference: emitters
produce bytes from the shared IR; renderers own graphics APIs. Types go
public as the backend's contract; AgcExports and the presenter consume
them exactly as before.

The ShaderDump tool drops its last reflection: with both halves of the
pipeline public it drives decode and all three emit entry points with
direct typed calls, retiring the PadWithDefaults invoke shim — and it
no longer references SharpEmu.Libs at all, making the conformance tool
emulator-independent by design. Verified by a clean solution build, the
test suite, a full ShaderDump conformance run, and a locked-mode
restore under the pinned SDK.

* [Gpu] Extract the guest-GPU backend seam (IGuestGpuBackend)

The AGC/VideoOut/SystemService export layers now reach the renderer
through IGuestGpuBackend via GuestGpu.Current (mirroring HostPlatform),
instead of calling VulkanVideoPresenter statics. The Vulkan backend is
a thin adapter over the existing presenter, so the extraction stays
mechanical; only the adapter and the presenter itself reference the
presenter now.

The types crossing the seam move to Gpu/GuestGpuTypes.cs and drop their
Vulkan prefixes, which an audit showed were misnomers: every field is a
neutral primitive or a raw guest value (guest addresses, format and
number-type codes, CB_BLEND register bitfields, verbatim sampler
descriptor dwords). The one genuine Vulkan value in the old surface —
the Silk.NET Format inside VulkanRenderTargetFormat, which callers
never read — stops crossing: TryDecodeRenderTargetFormat is replaced at
the seam by TryGetRenderTargetOutputKind, which surfaces only the
Gen5PixelOutputKind callers actually consume, keeping native formats a
backend-internal concern. ToVulkanSampler in AgcExports is renamed
ToGuestSampler to match what it always produced.

Seam rules are documented on the interface: no host-API value crosses,
and submission stays coarse-grained with synchronization internal to
backends. Interim exception, resolved next: shader parameters are still
SPIR-V blobs.

* [Gpu] Move shader compilation behind the guest-GPU backend

The seam's interim exception is gone: AgcExports no longer calls
Gen5SpirvTranslator or handles SPIR-V bytes. IGuestGpuBackend gains the
three TryCompile entry points, which take the backend-neutral
(Gen5ShaderState, Gen5ShaderEvaluation) contract plus the flat
per-role resource-slot bases a multi-stage draw needs, and return
opaque IGuestCompiledShader handles that only the producing backend can
submit — the Vulkan backend wraps its SPIR-V in
VulkanCompiledGuestShader and rejects foreign handles loudly. Draw and
dispatch submissions take handles instead of byte arrays; the shader
caches in AgcExports store handles.

IGuestCompiledShader.Payload exposes the backend-defined compiled bytes
for exactly two callers: the diagnostics dump and the size trace —
documented as never-interpret. The unused _pixelSpirvCache is deleted.
With this, a Metal or DX12 backend plugs in by implementing
IGuestGpuBackend with its own codegen; nothing in the export layers
knows which shader format exists.

Verified by a clean solution build, the test suite, and a full
ShaderDump conformance run under the pinned SDK.

* [Gpu] Fix rename collateral from the seam extraction

Address review findings: a doc comment picked up the mechanical
VulkanVideoPresenter -> GuestGpu.Current rewrite and ended up naming
members that do not exist on the interface, and CreateVulkanIndexBuffer
kept its Vulkan prefix while every sibling factory was de-Vulkanized —
it produces the neutral GuestIndexBuffer, so it is CreateGuestIndexBuffer.

* [Gpu] Label diagnostics dumps with the backend's payload extension

Address the review's altitude finding on DumpSpirv: the dump helper's
IR-disassembly half is backend-neutral and stays put, but writing the
opaque payload to a hardcoded .spv interpreted bytes the seam says
never to interpret. IGuestCompiledShader now declares its payload's
file extension, and the renamed DumpCompiledShader takes the handle and
writes honestly-labeled dumps whichever backend produced them.

* [Gpu] Make the shader-cache hit path allocation-free and lock-free

Every translated draw built its cache key with a LINQ Select feeding
string.Join plus one interpolated string per render target — steady
per-draw allocation whether or not the shaders were already cached. The
output layout is now packed exactly into a ulong (guest slot in 6 bits
+ output kind in 2 bits per target, host locations being the byte
positions, target count in the key beside it), and the
Gen5PixelOutputBinding array is only materialized on a cache miss,
where compilation dwarfs it.

The graphics/compute shader caches switch from Dictionary guarded by
_submitTraceGate to ConcurrentDictionary, making the per-draw and
per-dispatch hit paths lock-free and decoupling them from the tracing
gate they coincidentally shared. And the seam-shaped render-target list
is built once when a translated draw is created instead of a
Select/ToArray per submission of a cached draw.

* [Gpu] Replace LINQ with explicit loops in code this branch introduced

Project rule going forward: no LINQ — it allocates enumerators,
closures, and delegates, and this codebase is GC-pause-sensitive. The
pixel-output and guest-render-target array builds and the ShaderDump
store-PC collection become plain loops; pre-existing LINQ elsewhere is
left for changes that already touch those lines.

* [ShaderCompiler] Suppress CA2255 on the evaluator hook installer

The analyzer coverage that arrived with the rebase flags
ModuleInitializer in library code; this is the rule's intended advanced
scenario — the hook must be installed before any code path can reach
the evaluator, and every such path enters through this assembly — so
suppress with that justification rather than weaken the guarantee to a
static constructor's lazier timing.

* [Gpu] Resolve rebase artifacts onto main

Dedupe the System.Collections.Concurrent using in AgcExports that the
rebase merge duplicated (main and this branch each added it), and
regenerate the lock files for the new shader-compiler projects and
SharpEmu.Libs against main's current package graph so --locked-mode
restore matches at the branch tip.

* [CI] Comment per-platform build artifact links on PRs

Adds a workflow_run workflow that, after "Build and Release" finishes a
pull-request build, posts (and keeps updated in place) a single PR
comment linking the Windows, Linux, and macOS artifacts from that run.

It runs via workflow_run rather than in the build workflow because PRs
from forks build with a read-only token that cannot comment; the
follow-on run executes in the base-repo context with write access and
without checking out fork code. GitHub only triggers workflow_run from
the default branch, so this takes effect once merged to main.
2026-07-15 11:11:24 -06:00
kuba fa2616d224 Linux and macOS support (#47)
* [macos/linux] Cross-platform host memory, TLS, and ABI layer for POSIX

Introduces the foundation for running SharpEmu on macOS (osx-x64 under
Rosetta 2) and Linux (linux-x64). The CPU backend executes guest x86-64
code natively, so these targets run the whole process as x86-64; this
commit replaces the Windows-only host primitives with platform-dispatched
equivalents so the guest boots and services HLE calls off Windows.

Memory (HostMemory.cs, new): a Win32-semantics facade over
mmap/mprotect/munmap with a shadow region table answering VirtualQuery.
PhysicalVirtualMemory, DirectExecutionBackend, StubManager, and the two
Kernel*CompatExports now go through it instead of kernel32 P/Invokes.
Exact-address requests use MAP_FIXED_NOREPLACE (Linux) / guarded
MAP_FIXED (macOS) so they match Win32 "map there or fail" semantics.

TLS + host helpers (PosixHostStubs.cs, new): pthread-backed TLS and
Win64-ABI-compatible stubs for the kernel32 helpers the backend embeds
into emitted x86-64 code (TlsGetValue, QueryPerformanceCounter,
SwitchToThread, Sleep). A Win64->SysV thunk wraps managed callbacks,
since .NET on POSIX compiles them for the SysV ABI while the emitted
call sites use Win64.

Guest address layout: the 0x7FFx window is Windows-only (dyld shared
cache / Rosetta runtime live there on POSIX), so stack/TLS/stub regions
relocate to 0x6FFx off Windows.

Vectored exception handling is gated off on POSIX for now (guest faults
are not yet recovered) — the signal-based bridge is the next step. Also
adds osx-x64 to the RID list and a Docker-based Linux smoke-test script.

Status: on both macOS (Rosetta) and Linux (amd64), the guest now boots,
runs native x86-64 code, and dispatches HLE imports. macOS stops at a
Rosetta translation-cache issue; Linux runs ~252 imports through C++
static-init before hitting the missing fault handler (SIGSEGV).

* [posix] Bridge the vectored exception handler to sigaction(SIGSEGV/SIGBUS/SIGILL)

Guest faults on macOS/Linux previously terminated the process because the
recovery logic in DirectExecutionBackend.Exceptions.cs was Windows-only.
This adds a POSIX front-end that reuses the existing handler bodies:

- DirectExecutionBackend.PosixSignals.cs installs SA_SIGINFO handlers via
  an [UnmanagedCallersOnly] entry, rebuilds the Win64 EXCEPTION_POINTERS /
  CONTEXT view from the platform mcontext (Darwin __ss thread state via
  the mcontext pointer at ucontext+48, Linux glibc gregs at ucontext+40 --
  offsets verified against the headers on both platforms), runs the same
  chain as the VEH path (TryRecoverUnresolvedSentinel trap-sentinel
  recovery, TryHandleLazyCommittedPage demand paging, VectoredHandler
  diagnostics incl. FS/GS TLS-fault detection), and writes register
  changes back into the mcontext so sigreturn resumes the repaired guest.
  Unrecovered faults chain to the previously installed handler so the
  .NET runtime keeps mapping its own faults to managed exceptions.

- The whole recovery path is warmed up with fabricated inputs before the
  handlers are installed. This is required under Rosetta 2: the signal
  trampoline cannot enter x86 code that has never been executed (and so
  never translated) -- a cold handler is silently never invoked and the
  faulting instruction retries forever (reproduced and verified in an
  isolated .NET test under Rosetta for Linux). It also keeps first-fault
  JIT work out of the signal frame.

- Handlers run without SA_ONSTACK: the runtime's alternate stacks are too
  small for the diagnostic path, while guest (2MB) and host thread stacks
  match where Windows dispatches exceptions anyway.

- The raw reads in the shared fault diagnostics (stack qwords, RBP walk,
  code bytes at RIP) now probe the region table on POSIX before touching
  memory, since a nested SIGSEGV inside the handler would kill the
  process before diagnostics finish. Windows keeps its try/catch reads.

- Escape hatches: SHARPEMU_DISABLE_POSIX_SIGNALS=1 skips installation,
  SHARPEMU_DISABLE_RAW_HANDLER=1 disables sentinel recovery (parity with
  Windows), SHARPEMU_LOG_POSIX_SIGNALS=1 traces every delivery (first 16
  and every 1024th are always traced).

Verified with the test game: Linux (amd64 container) previously died with
SIGSEGV right after import #252; it now recovers/diagnoses signals and the
run proceeds to the real next blocker, an unpatched negative-offset guest
TLS read (fault at TLS base - 0x1708), which gets the full NATIVE
EXCEPTION dump before terminating. macOS is unchanged: the bridge installs
and the game still stops at the known Rosetta translation-cache error at
import 12, which is the next work item.

* [posix] Fix guest memory layout faults: TLS prefix, exact mmap, map search base

Three fixes that take the test game from dying during libc init to running
its full main loop on macOS and Linux:

- Static TLS blocks live below the TCB (FreeBSD amd64 variant II) and
  libc.prx reaches past -0x1700, but only a 4KB prefix was mapped below
  the TLS base. The prefix is now 64KB on POSIX (Windows keeps 4KB); the
  fault was a read at TLS base - 0x1708 during libc init.

- HostMemory exact allocation on macOS used MAP_FIXED, which silently
  maps over untracked host memory. The direct-memory allocator's address
  scan walked into the .NET runtime's JIT heap and replaced live code,
  which under Rosetta 2 surfaced as "no code fragment associated with
  the given arm pc". Exact placement now passes the address as a hint
  and fails on relocation, like MAP_FIXED_NOREPLACE does on Linux.

- sceKernelMapDirectMemory/MapFlexibleMemory searched for free space
  starting at 4GB, which is the Mach-O image base on macOS. The default
  search base is 0x20_0000_0000 on POSIX, and TryAllocateAtOrAbove now
  asks the kernel for a placement instead of page-stepping through host-
  owned address space (Rosetta ignores mmap hints for whole VA windows),
  over-allocating when the caller needs more than page alignment.

Windows behavior is unchanged; all divergences are platform-guarded.

* [macos] Video presenter on the main thread, MoltenVK support, window keyboard input

Gets the test game from a headless loop to a playable window on macOS:

- AppKit traps with SIGILL ("NSUpdateCycleInitialize() is called off the
  main thread") when GLFW runs on a worker thread. The CLI now moves
  emulation onto a worker thread on macOS and parks the real main thread
  in HostMainThread.Pump(); the presenter posts its whole window loop
  there instead of spawning a thread, and a shutdown handler asks the
  render loop to close the window so the pump unwinds on guest exit.

- MoltenVK: enable VK_KHR_portability_enumeration (+ the portability
  instance flag) and VK_KHR_portability_subset when advertised, and gate
  robustBufferAccess2 on the device actually supporting it (Metal does
  not; the old code keyed it off robustImageAccess2 and vkCreateDevice
  failed with ErrorFeatureNotPresent).

- Input: pad exports polled user32 GetAsyncKeyState, so POSIX hosts threw
  DllNotFoundException per scePadReadState call. The presenter now
  attaches the window's keyboard via Silk.NET.Input into HostWindowInput,
  and the pad exports map the existing VK-code layout onto it off
  Windows. Headless hosts (Linux containers) report a disconnected
  keyboard and fall back to neutral pad data silently.

GLFW needs an x86-64 Vulkan loader under Rosetta: place a universal
libMoltenVK.dylib next to SharpEmu named libvulkan.1.dylib (Homebrew's
arm64-only copy cannot load into the x86-64 process) and export
DYLD_LIBRARY_PATH to that directory.

Verified: Dreaming Sarah boots to a MoltenVK-backed 2560x1440 window on
macOS (Apple M4, Rosetta 2), renders the intro, title, and menus, and
keyboard input drives it into gameplay. Linux (amd64 container) runs the
same build headless through millions of imports with no faults. Windows
paths unchanged; arm64 and x64 builds clean.

* [posix] CoreAudio playback, self-contained MoltenVK loading, input/log polish

- Audio: sceAudioOut ports now play through an AudioQueue backend on macOS
  (stereo PCM16 with the same 32KB backpressure pacing as the WinMM path).
  The WinMM port and the new CoreAudio port share an IHostAudioPort
  interface and sample converter; hosts without a backend (Linux
  containers) keep the silent fallback.

- MoltenVK: GLFW resolves Vulkan with dlopen("libvulkan.1.dylib"), which
  cannot see the app-local universal MoltenVK build, so the presenter now
  feeds vkGetInstanceProcAddr straight into glfwInitVulkanLoader (GLFW
  3.4) before creating the window. No DYLD_LIBRARY_PATH needed; the CLI
  also preloads the dylib for Silk.NET and prints setup hints when it is
  missing. scripts/fetch-macos-moltenvk.sh stages the official universal
  dylib next to a build.

- The virtual-range allocator's failure trace now names the address and
  length instead of "AllocateAt invocation threw".

Investigated and documented (not port defects): the savedata transaction
failure is identical on Linux and macOS (HLE argument-register mapping for
sceSaveDataCreateTransactionResource), and the in-game tile speckling has
no platform-specific code in its path - the one macOS-only delta is that
MoltenVK lacks robustBufferAccess2, so out-of-bounds shader reads return
garbage instead of zeros.

Verified on macOS: window, audio backend, and keyboard input all come up
with zero environment configuration; the game runs to gameplay. Linux
headless run unchanged (silent audio, no faults). Windows paths untouched;
arm64 and x64 builds clean.

* [cpu] Preserve guest registers and flags across patched TLS accesses

The TLS patch handler replaces guest `mov reg, fs:[...]` instructions,
which preserve every other register and the flags - but the handler
loaded the TLS index into ecx and called TlsGetValue (Win64: clobbers
rcx/rdx/r8-r11) with `sub/add rsp` trashing the arithmetic flags. Guest
code that keeps live values or comparison results across a TLS access
computed garbage deterministically. The handler now saves rcx, rdx,
r8-r11, and the flags around the call, keeping the same inner stack
alignment. This applies to the load patches and both store-helper stubs,
on every platform.

Also in this change, from the rendering-artifact investigation:

- The present blit picks linear filtering for any fractional scale
  (nearest only for integer upscales): a 3840x2160 guest frame blitted
  into a 2560x1440 swapchain with nearest silently dropped every third
  row/column.
- ClampViewport no longer trims the guest viewport rectangle to the
  render target; trimming changed the guest's scale/offset and skewed
  texel addressing. Vulkan permits viewports beyond the framebuffer
  (the scissor confines rendering), so only spec bounds are enforced.
- Env-gated diagnostics grown during the investigation: guest texture
  dumps (SHARPEMU_TEXTURE_DUMP_DIR), aliased guest-image readback dumps
  (SHARPEMU_TRACE_GUEST_IMAGES=alias), small-render-target write movies
  (SHARPEMU_TRACE_GUEST_WRITES=small), unattended input injection
  (SHARPEMU_AUTO_CROSS=secs,...), viewport nudging
  (SHARPEMU_VIEWPORT_EPSILON), chunked-draw toggle
  (SHARPEMU_DISABLE_CHUNKED_DRAWS), and rect-list/draw vertex traces.

Known remaining issue (root cause narrowed, not yet fixed): the game's
terrain texture pages are corrupted in guest memory before any GPU work
- the mound's solid-fill 32x32 tiles decode to fully transparent texels
and the grass page has deterministic gaps, byte-identical across runs.
Ruled out: memcpy/memmove/memset/realloc HLE semantics, sampler wrap
modes, texel-boundary rounding, chunked draws, viewport handling. Next
step is auditing the Chowdren asset decode path (custom compressed
images) against the emulator's import surface.

* [linux] ALSA playback backend for sceAudioOut

sceAudioOut ports on Linux now play through libasound instead of the
silent fallback. The PCM device opens in blocking mode with ~170ms of
device buffer (the time-equivalent of the 32KB queue the WinMM and
CoreAudio ports keep), so snd_pcm_writei provides the same backpressure
pacing without a managed queue. Underruns and suspend/resume go through
snd_pcm_recover with one retry per submit; anything else drops the
buffer rather than stalling the guest.

The "default" device routes through PulseAudio/PipeWire on desktops
and straight to hardware on bare ALSA; SHARPEMU_ALSA_DEVICE overrides
it (the null device makes the path testable in containers). A missing
libasound or device fails port creation and lands in the existing
silent fallback.

Verified in an amd64 container: the test game opens the port
(backend=alsa, 48kHz stereo float32) and streams sceAudioOutOutput
through the null device for a full run; without a usable device the
port logs a warning and falls back to silent. Playback on real Linux
audio hardware has not been tested.

* [fixes] Address review feedback: commit bounds, CoreAudio shutdown, dump errors

- HostMemory: a MEM_COMMIT that runs past its reservation now fails like
  Win32 instead of committing a prefix and reporting success. All current
  callers already clamp their ranges to the region, so this only guards
  future callers.

- CoreAudioPort: Dispose wakes a submitter waiting on backpressure and
  the wait treats ObjectDisposedException as a timed-out wait, so closing
  a port during playback can no longer throw. A failed AudioQueueStart
  tears the queue down and fails fast instead of leaving an undrainable
  queue that stalls every later submit on its timeout.

- AgcExports: texture dumping catches all write failures (bad path,
  permissions), logging a warning instead of crashing when
  SHARPEMU_TEXTURE_DUMP_DIR points somewhere unusable.

Verified with the Linux container run: game boots and streams audio with
the stricter commit check, and a dump dir under /proc produces warnings
instead of taking the process down.

* [ci] Build linux-x64 and osx-x64 archives

Adds a build-posix matrix job (ubuntu-latest / macos-latest) mirroring
the Windows build: locked restore, Release build, self-contained CLI
publish, and a tar.gz artifact per RID (tar keeps the executable bit).
The macOS archive also stages the universal MoltenVK dylib via
scripts/fetch-macos-moltenvk.sh so the artifact runs without any manual
Vulkan setup. The release job still only ships the Windows archive.

* [cli] Keep POSIX glfw natives outside the single-file bundle

The KeepGlfwOutsideSingleFile target only matched filenames starting
with 'glfw', which covers Windows (glfw3.dll) but not libglfw.3.dylib /
libglfw.so.3. Those got embedded into the single-file bundle, and
Silk.NET's library loader does not probe the bundle extraction
directory, so a published build died with "Couldn't find a suitable
window platform" (and the glfwInitVulkanLoader wiring, which loads the
library from AppContext.BaseDirectory, could not run either). Keeping
the POSIX names loose next to the executable fixes both, the same way
the Windows build already handled it.

Found by running the CI-built osx-x64 archive: video failed while local
loose-file builds worked. With the fix the published single-file build
opens the MoltenVK window, wires the loader, and reaches gameplay.

* [ci] Publish linux-x64 and osx-x64 release archives

The build-posix artifacts now ship as per-RID GitHub releases on main
pushes and manual dispatches, tagged the same way as the win64 ones
(<rid>-<ref>-<sha>). Archives stay tar.gz so the executable bit
survives extraction.

* [cli] Fail early on non-x86-64 host processes

The CPU backend executes guest x86-64 code natively, so the process
must be x86-64 (win-x64/linux-x64 on x64 hardware, osx-x64 under
Rosetta 2 on Apple Silicon). An arm64 process previously failed deep
inside emulation startup, indistinguishable from MoltenVK, signal
handler, or guest memory problems. CLI mode now checks the process
architecture up front and exits with a message naming the supported
execution model (and the Rosetta install command on macOS). The
GUI-only path stays usable on arm64.

* [video] Log the selected Vulkan device name and API version

The presenter never named the GPU it picked, so a 'no video' report
could not be told apart from a real windowing failure without guessing.
It now logs the device name, type, and API version right after
selection. A software rasterizer (llvmpipe/lavapipe/SwiftShader) shows
up here and typically lacks the device features the translated shaders
need, which is the likely cause when a window opens and presents frames
but nothing draws.

* [video] Steer GLFW to XWayland on Wayland sessions

GLFW's native Wayland backend does not reliably map the Vulkan window
with some drivers (NVIDIA in particular): frames present but the window
never becomes visible, so the game runs with audio and no picture. A
report on an RTX 5080 showed exactly this — all device features present,
frames presenting, but the log had 'libdecor-gtk.so failed to init' and
a 1.4x-scaled window, both Wayland tells.

On a Wayland session that also exposes an X server (DISPLAY set), the
presenter now clears WAYLAND_DISPLAY for its own process before GLFW
initializes, so GLFW selects its dependable X11/XWayland backend.
SHARPEMU_ENABLE_WAYLAND=1 opts back into native Wayland. Headless
(no DISPLAY) and non-Linux hosts are unaffected.

* [video] Force GLFW X11 backend via the platform init hint, log the platform

The previous attempt cleared WAYLAND_DISPLAY to steer GLFW off Wayland,
but a reporter still hit the native-Wayland path (the Wayland-only
libdecor error persisted), so that env trick doesn't switch GLFW.

Use GLFW's supported mechanism instead: glfwInitHint(GLFW_PLATFORM,
GLFW_PLATFORM_X11) before GLFW initializes, called into the same libglfw
GLFW itself loads (the pattern InitializeMacVulkanLoader already uses).
Still gated on a Wayland session with an X server present (DISPLAY set)
so we never force X11 where XWayland can't catch it, and still
overridable with SHARPEMU_ENABLE_WAYLAND=1.

Also logs 'GLFW windowing platform in use: <backend>' after init via
glfwGetPlatform, so a 'no window' report shows X11 vs Wayland outright.
Verified on macOS: the readback correctly reports Cocoa and the
presenter is unaffected (the fix is a no-op off Linux).

* [video] Run the GLFW window on the main thread on Linux too

GLFW requires window creation and event processing on the process main
thread on every platform: initialization, window creation, and
glfwPollEvents are main-thread-only, and X11 in particular has a single
event queue that must be serviced there. A window created and polled on
another thread may never map — which is why the game ran (audio, imports,
even Vulkan present) with no visible window on Linux.

macOS already routed the window loop to the main-thread pump (AppKit
needs it); Windows is fine because it has a per-thread event queue. Linux
was the gap: it spawned a background thread for the presenter. Extend the
existing HostMainThread pattern to Linux — emulation runs on a worker,
the main thread pumps the window work the presenter posts.

Refs GLFW intro guide (thread-safety): init, window creation, and event
processing are restricted to the main thread.

Verified: macOS still boots to its window unchanged; the Linux headless
container runs to millions of imports with no deadlock or regression.
On-screen confirmation on a real Linux desktop is still pending, but this
is the documented root cause for a windowless-but-running Linux session.

* [posix] Skip Win32 native guest workers

* [vulkan] Synchronize offscreen targets before present

* [vulkan] Transition fresh textures from undefined layout

* [vulkan] Report swapchain pixels before source readback

* [vulkan] Emit requested guest image diagnostics

* [agc] Diagnose guest texture fallbacks

* [linux] Keep guest GPU mappings in low address space

* [video] Reduce diagnostic stalls and drain complete frames

* [memory] Harden packed GPU address handling

* [readme] Document Linux and macOS release support

* [posix] Integrate the host platform abstraction

* [posix] Restore guest thread address window

* [video] Run the performance HUD on POSIX hosts

The FPS/CPU/work HUD bailed out unless the host was Windows; only the
per-thread hottest-thread scan actually needs Windows APIs. Keep that
scan Windows-only (POSIX reports 'idle') and let the rest of the HUD
run everywhere — the title is already set from the render thread, which
owns the window on macOS and Linux.

* [posix] Implement native guest worker threads

Guest entry stubs must not run above CLR-managed frames on CLR-created
threads (see the NativeWorker preamble); the PR previously fell back to
the inline calli path on POSIX, which reproduced the documented
'attempted to call a UnmanagedCallersOnly method from managed code'
fail-fast (observed after Dreaming Sarah's menu select) and left the
runtime's suspension machinery walking guest frames.

Provide the missing POSIX half of the worker loop:
- PosixHostStubs grows Win64-convention WaitForSingleObject/SetEvent/
  ExitThread stubs backed by dispatch semaphores (macOS) / unnamed POSIX
  semaphores (Linux) plus pthread_exit, with EINTR retry in the wait.
- Worker events are creatable/signalable/waitable from managed code too,
  so NativeGuestExecutor.Run keeps its handshake (AutoResetEvent stays
  on Windows byte-for-byte).
- PosixHostThreading implements CreateNativeThread/WaitForThreadExit/
  CloseThreadHandle over pthreads (liveness probed with
  pthread_kill(0), then joined).
- RunPrologue/RunEpilogue are routed through the existing Win64->SysV
  thunks, so the emitted loop stays identical across platforms.

* [macos] Disable concurrent GC under Rosetta's write-watch hazard

Background GC's write-watch revisit (SoftwareWriteWatch::GetDirty ->
FlushProcessWriteBuffers) calls thread_get_register_pointer_values on
every thread; under Rosetta 2 that Mach call stalls indefinitely on
threads executing translated guest code. The background mark phase then
never finishes and every allocating or Monitor-taking thread wedges
behind it — observed as Dreaming Sarah freezing at the menu/loading
screen with FPS 0 in 5 of 7 runs, dispatcher/watchdog parked in
Monitor.Enter and all BGC threads waiting in t_join.

Non-concurrent GC never takes that path; a 5-minute soak now holds
22-31 fps in-game with zero stalls. Windows and Linux keep concurrent
GC.

* [diag] Periodic guest-thread snapshots with gate-owner tracking

SHARPEMU_PERIODIC_SNAPSHOT_SECONDS=N dumps the stall snapshot every N
seconds even while imports are progressing, for soft stalls where the
game stops advancing but threads keep spinning. The periodic dump never
touches the guest-thread gate (it must keep reporting when the gate is
what's wedged): it reads a lock-free owner record — every gate
acquisition now goes through LockGate(site), which notes site/thread —
and walks the thread table without the lock, tolerating torn reads.
SHARPEMU_PERIODIC_SNAPSHOT_FILE redirects the dump to a side file for
the case where the console itself is wedged (frozen log mirror was one
of the observed failure modes).

* [nuget] Add osx-x64 RID targets to lock files

* [cpu] Back off the guest join poll

TryJoinThread polled the host thread at a fixed 1ms; a game main thread
joining a long-lived worker (Dreaming Sarah parks there for the whole
session) burned ~5% of managed CPU in Join/Sleep syscalls. Ramp the
poll interval to 10ms once the join is clearly long-lived — exit
detection latency for long joins moves from ~1ms to at most 10ms, and
short-lived joins still resolve on the first 1ms polls.

* [nuget] Add linux-x64/win-x64 RID targets to lock files

* [posix] Keep guest stacks clear of the import-stub descent

The import-stub region descends from 0x7000_0000_0000 on the same 16MB
grid as the guest thread windows; moving stacks to 0x6FFF_E000_0000 put
them inside the stub region's 64-module descent range (floor
0x6FFF_C000_0000), silently consuming the top ~32 stack slots on hosts
with many loaded modules. Drop the POSIX stack base to 0x6FFF_A000_0000:
512MB below the stub floor, still 2.5GB above the TLS window. Windows
keeps 0x7FFF_E000_0000 (its bands are ~15TB apart).

* [pad] Read window gamepads on POSIX hosts

XInput and the DualSense hid reader are Windows-only, which left
macOS/Linux with keyboard input only. The presenter's Silk/GLFW input
context already enumerates gamepads on both platforms, so track their
state in HostWindowInput (event-driven on the window thread, snapshot
guarded like the key set) translated to ORBIS conventions: GLFW's Xbox
layout maps A/B/X/Y to Cross/Circle/Square/Triangle, sticks bias from
-1..1 to 0..255 with Y growing down, and triggers rescale from GLFW's
-1..1 resting-at--1 range with digital L2/R2 bits past 25%.

The merge into ReadHostInputState is gated to non-Windows so a physical
pad is never sampled twice through both a native reader and GLFW.
Hotplug is handled via ConnectionChanged; with no pad connected the
path is inert.

Untested against a physical controller (none attached to the dev host);
axis conventions follow the GLFW gamepad-mapping contract.

* [nuget] Refresh lock files after cross-RID restores

* [posix] Adopt the host audio/input seams from main

Main's #192 abstracted audio output and pad/keyboard input behind
IHostAudioOutput/IHostInput; re-express the POSIX backends behind them:

- CoreAudioPort/AlsaAudioPort move to Host/Posix as
  PosixCoreAudioStream/PosixAlsaAudioStream implementing
  IHostAudioStream. The seam converts to stereo PCM16 before Submit, so
  the ports' own conversion (and IHostAudioPort/AudioSampleConverter)
  is gone; queueing and backpressure are unchanged.
- PosixHostAudio selects CoreAudio (macOS) / ALSA (Linux) as the
  platform's IHostAudioOutput.
- PosixHostInput implements IHostInput over an
  IPosixWindowInputSource that HostWindowInput registers when the
  presenter attaches the window's GLFW input context: keyboard with
  virtual-key translation, the window gamepad snapshot (now in seam
  HostGamepadState/HostGamepadButtons terms), and keyboard-connected as
  the focus signal. Rumble/lightbar no-op (GLFW has no such API).
- PadExports drops its direct HostWindowInput gamepad merge — pads now
  flow through IHostInput.GetGamepadStates like every platform.
- PosixHostThreading.RequestTimerResolution is a documented no-op.

All three RIDs build; SharpEmu.Libs.Tests pass (26/26).

* [nuget] Regenerate GUI lock file for RID-less locked restore

Local cross-RID builds stamped a win-x64 runtimes section into
SharpEmu.GUI's lock file; the project declares no RuntimeIdentifiers,
so CI's 'dotnet restore --locked-mode' failed with NU1004 on every
platform. Regenerated via a plain solution restore (--force-evaluate),
matching what the workflow validates.
2026-07-15 15:36:20 +03:00
Gutemberg Ribeiro 62e1775c5c [HLE] Remove steady-state allocations from the hot HLE paths (#190)
* [HLE] Stop allocating on the memcpy/memset and trace hot paths

memcpy/memmove no longer allocate a bounce buffer sized to the whole
copy (large copies previously landed on the LOH); they loop through a
single pooled 256 KB rental, copying high-to-low when the destination
overlaps above the source so memmove semantics survive the chunking.
memset reuses a shared zero chunk for the dominant zero-fill case and
rents/fills only min(length, 16K) bytes for non-zero values instead of
allocating and filling a fresh 16 KB array per call; the map-time
zero-fill loop shares the same zero chunk.

SHARPEMU_LOG_SEMA / SHARPEMU_LOG_VIDEOOUT are now read once into cached
bools and every TraceSemaphore/TraceVideoOut call site is guarded, so
trace messages are no longer interpolated (and the env var no longer
queried) on every semaphore op and every flip with tracing off. Trace
output when the flags are set is unchanged.

* [HLE] Remove per-frame allocations from the vblank/flip/equeue plumbing

The 60 Hz vblank pump no longer allocates per edge: PumpVblanks reuses a
pump-thread-only port list instead of a LINQ Where/ToArray, and
SignalVblank/SubmitFlip snapshot their event registrations into pooled
rentals instead of copying the List on every edge and every flip (the
snapshot must still be taken, since triggers run outside _stateGate and
a per-port reusable buffer would race the pump thread against a guest
thread's first-edge signal).

sceKernelWaitEqueue delivery rents the dequeue buffer from the pool
instead of allocating an array per wait, and event-queue wake keys are
formatted once per handle (cached in a ConcurrentDictionary, dropped on
queue delete) instead of building the string on every enqueue. The
semaphore wake key moves onto KernelSemaphoreState at creation, the
same pattern the pthread mutex state already uses, removing the
per-signal/per-wait formatting. SHARPEMU_LOG_EQUEUE is read once into a
cached bool like the sema/videoout flags.

* [HLE] Read guest C-strings without per-call buffer allocations

CpuContext.TryReadNullTerminatedUtf8 allocated a byte[capacity] and
issued one TryRead per byte for every string-argument import. It now
reads through a stack buffer (pooled above 512 bytes) in 128-byte bulk
chunks, falling back to per-byte reads only when a chunk touches an
unreadable range so a terminator sitting just before unmapped memory
still resolves exactly as before. The chunk bound also keeps the
overread past the terminator smaller than the old loop's worst case is
wide, so no fault can appear where the byte loop succeeded.

TryReadAsciiZ (dlsym/symbol resolution) drops its List<byte> + ToArray
round-trip for the same stack/pooled buffer, keeping the byte-by-byte
TryReadByteCompat reads because their Marshal.ReadByte fallback must
probe exactly up to the terminator. Only the final string is allocated
on either path now.

* [HLE] Replace blocking-wait closures with waiter continuation objects

Every wait that actually parked a guest thread allocated two capturing
lambdas (plus their display classes) for the scheduler's resume/wake
callbacks. RequestCurrentThreadBlock and the backend's blocked-thread
state now carry a single IGuestThreadBlockWaiter instead of the
Func<int>/Func<bool> pair: TryWake keeps the run-under-the-scheduler-
gate contract and Resume still produces the guest's RAX on the woken
thread. The waiter stays attached through the wake transition (the old
code nulled only the wake handler there) and is consumed at resume.

The existing waiter objects absorb the captured state as fields, so a
blocking wait now allocates exactly one object: SemaphoreWaiter,
PthreadMutexWaiter, and EventFlagWaiter implement the interface
directly, and the equeue, cond, and rwlock waits get small waiter
classes replacing their closures. Handler bodies delegate to the same
static methods with the same arguments as before; the untimed event
flag wait's mutable captured result becomes a field on its waiter.

* [HLE] Back pending event queues with a ring deque instead of LinkedList

LinkedList<KernelQueuedEvent> allocated a node object on every
non-coalesced enqueue — one per vblank/flip edge per registered queue,
60+ times a second in steady state. KernelEventDeque is a grow-only
ring buffer over a KernelQueuedEvent[] with the three operations the
queue actually uses (AddLast, RemoveFirst, find-and-update-in-place by
ident/filter), so steady-state enqueue/dequeue allocates nothing and
the coalescing update writes the struct back through an indexer instead
of a node reference. All accesses stay under _eventQueueGate, matching
the LinkedList usage it replaces.

* [HLE] Cap memcpy chunk iterations at the requested size, not the rented length

Address Copilot review: ArrayPool.Rent may return a larger array than
requested, so sizing each iteration by chunk.Length let the copy
granularity depend on pool bucketing internals instead of the intended
256 KB chunking. Behavior was already correct for any chunk size (each
iteration re-reads the source, and the overlap ordering is size-
independent), but the loop now mins against the requested chunkLength,
matching what memset already does.

* [HLE] Skip the flip/vblank snapshot rental when no events are registered

Address Copilot review: SignalVblank and SubmitFlip rented (and
returned) a pooled snapshot even with zero registrations — steady
per-frame pool traffic for games that never register flip events and
only poll flip status. Zero-count signals now skip the rental, the
copy, and the trigger loop entirely, which also retires the
Math.Max(count, 1) minimum-rent guard.
2026-07-15 12:57:40 +03:00
Spooks 9d88542efd Fix virtual memory allocation and access (#193)
* Fix virtual memory allocation and access

* Update test dependency lock file
2026-07-14 21:50:54 -06:00
Gutemberg Ribeiro f23161be9a Host platform abstraction layer for the execution engine (#181)
* [Host] Introduce host platform abstraction with IHostMemory

Add SharpEmu.HLE/Host with IHostPlatform/IHostMemory interfaces, neutral
page-protection/region enums, and a HostPlatform.Current factory that
resolves the Windows backend (or throws PlatformNotSupportedException on
other OSes, matching today's de-facto behavior). WindowsHostMemory wraps
the exact VirtualAlloc/VirtualFree/VirtualProtect/VirtualQuery calls used
across the engine today, with identical MEM_*/PAGE_* constants.

Migrate StubManager as the first consumer: its private kernel32 P/Invokes
and enums are replaced by IHostMemory calls that issue the same two
native operations (RWX commit+reserve of the PLT arena, release on
Dispose). No behavior change.

This is the first step toward supporting non-Windows hosts; subsequent
commits move the remaining direct P/Invokes in Core and Libs behind the
same seam.

* [Host] Route PhysicalVirtualMemory through IHostMemory

Replace the class's private VirtualAlloc/VirtualFree/VirtualProtect/
VirtualQuery P/Invokes with IHostMemory calls. Every site maps 1:1 onto
the exact native call it issued before: MEM_COMMIT|MEM_RESERVE ->
Allocate, MEM_RESERVE -> Reserve, fault-path commits -> Commit, and
MEM_RELEASE -> Free, with identical protection values produced by the
Windows backend.

IHostMemory gains ProtectRaw so the save/restore protection sequences in
TryWriteExclusive and TryTemporarilyProtectForRead round-trip the raw OS
protection word (including modifier bits the neutral enum cannot
represent) exactly as before. Raw PAGE_* constants remain only for the
internal region-classification helpers, which only ever see values this
class itself assigned.

The exact-address free-on-mismatch, lazy reserve-only threshold, prime
loop, and all trace strings are unchanged.

* [Host] Add IGuestAddressSpace and retire the reflection-based allocator lookup

Introduce IGuestAddressSpace in SharpEmu.HLE (fixed-address AllocateAt /
TryAllocateAtOrAbove and guest mprotect via TryProtect) with signatures
copied from PhysicalVirtualMemory, which now implements it. TryProtect
reproduces the read/write/execute decomposition that
KernelMemoryCompatExports.ResolveHostProtection performs, yielding the
same PAGE_* values through the Windows backend.

KernelVirtualRangeAllocator previously located AllocateAt via cached
MethodInfo reflection (because SharpEmu.Libs cannot see Core types) and
walked wrapper memories through an untyped 'Inner' property. Both are
now typed: ICpuMemoryWrapper exposes the decorated memory (implemented
by TrackedCpuMemory, whose Inner property already existed) and the
allocator type-tests for IGuestAddressSpace with the same bounded
unwrap depth. Failure paths keep the exact [LOADER][TRACE] strings.

* [Host] Move Kernel HLE memory exports off direct kernel32 P/Invokes

KernelMemoryCompatExports loses its private VirtualQuery/VirtualProtect/
VirtualAlloc/VirtualFree declarations and MemoryBasicInformation struct:

- Guest mprotect (sceKernelMprotect/sceKernelMtypeprotect) now routes
  through IGuestAddressSpace.TryProtect resolved from ctx.Memory. The
  orbis read/write/execute decomposition moves into a GuestPageProtection
  conversion whose mapping is value-identical to the removed
  ResolveHostProtection.
- The guarded libc heap and host-page accessibility checks go through
  IHostMemory (same commit+reserve/protect/free sequence; guard-page and
  protection-mask checks compare HostRegionInfo.RawProtection against the
  same PAGE_* literals as before).
- HostMemory is exposed as a property so merely loading the type never
  resolves the platform backend on non-Windows hosts.

KernelRuntimeCompatExports' RDTSC stub allocates its 16-byte RWX page via
IHostMemory.Allocate; the OperatingSystem.IsWindows() gate returning null
is unchanged.

* [Host] Abstract thread, TLS, and symbol primitives in the execution backend

Add IHostThreading (native TLS slots, current-thread id, affinity, raw
thread create/join, diagnostic register capture) and IHostSymbolResolver
(enum-keyed host function addresses baked into emitted stubs), with
Windows implementations wrapping the exact kernel32 calls the backend
made directly before.

DirectExecutionBackend takes an optional IHostPlatform (defaulting to
HostPlatform.Current) and routes every TlsAlloc/TlsFree/TlsSet/GetValue,
GetCurrentThreadId, SetThreadAffinityMask, GetModuleHandle/GetProcAddress
and the suspend+GetThreadContext diagnostic snapshot through it. The
snapshot moves wholesale into WindowsHostThreading (including the Win64
CONTEXT size/flags/offsets, which are Windows-specific by nature) and
returns a neutral HostCapturedRegisters.

NativeGuestExecutor resolves WaitForSingleObject/SetEvent/ExitThread via
the symbol resolver — the same addresses end up in the emitted run loop,
so stub bytes are unchanged — and creates/joins its raw worker thread
through IHostThreading with the same stack-reservation semantics. The
run-loop emitter itself does not move.

Marshal.GetLastWin32Error() in the affinity-failure log still observes
SetThreadAffinityMask's error because the wrapper makes no intervening
SetLastError call.

* [Host] Move fault handling and remaining backend memory ops behind the seam

Add IHostFaultHandling (handler-thunk creation, first-chance handler
install/remove, unhandled-filter set) with WindowsFaultHandling in a new
Cpu/Native/Windows/ folder. The exception-handler trampoline emitter
moves there whole — same pre-filtered NTSTATUS codes, same TEB gs:[8]/
gs:[0x10] stack-limit reads, same host-RSP TLS switch — parameterized
only by (managed callback, TLS slot, TlsGetValue address), which is
exactly what SetupExceptionHandler passed it before. Handler
installation order, the AddVectoredExceptionHandler(first=1) flag, the
SHARPEMU_DISABLE_RAW_HANDLER gate, and all install/teardown log strings
are unchanged.

Every remaining VirtualAlloc/VirtualProtect/VirtualFree/VirtualQuery/
FlushInstructionCache in the backend partials routes through IHostMemory
with 1:1 call mapping (RWX emit -> RX downgrade -> flush for stub
emission, reserve/commit for the PRT aperture and lazy-commit fault
path, raw-protection round-trips via ProtectRaw). HostRegionInfo gains
RawState/RawAllocationProtection so the lazy-commit trace lines and
protection-mask checks keep printing and comparing the exact native
values.

Windows semantics leaked as bare literals become named constants with
identical values: NTSTATUS codes (WindowsFaultCodes) and Win64 CONTEXT
byte offsets (Win64ContextOffsets, with the existing CTX_* constants
aliased to it and handler-local numeric offsets replaced by the names).

* [Host] Resolve the host platform explicitly at the composition root

SharpEmuRuntime.CreateDefault() now resolves HostPlatform.Current once
and passes it explicitly to PhysicalVirtualMemory and (via a new
optional CpuDispatcher parameter) to DirectExecutionBackend, replacing
the implicit default-argument fallbacks. On unsupported OSes boot now
fails at the root with PlatformNotSupportedException and a clear
message instead of on the first native call. A future Linux/macOS
backend plugs in by returning a different IHostPlatform here.

* [Host] Convert the platform backends to source-generated P/Invokes

Replace [DllImport] with [LibraryImport] in the four Windows backend
files added by this branch (WindowsHostMemory, WindowsHostThreading,
WindowsHostSymbolResolver, WindowsFaultHandling). Marshalling stubs are
now generated at compile time instead of JIT-emitted at runtime, which
fits the pre-JIT-everything boot model and keeps the backends
NativeAOT/trimming ready.

Interop stays zero-copy: all signatures are blittable, GetModuleHandleW
now pins the managed string via Utf16 marshalling instead of copying,
and GetProcAddress names marshal through a stack-allocated Utf8 buffer.
Implicit contracts become explicit where LibraryImport requires it:
TlsFree/TlsSetValue gain [MarshalAs(UnmanagedType.Bool)] (the 4-byte
Win32 BOOL DllImport assumed silently), and GetModuleHandle targets the
W entry point directly since LibraryImport never probes suffixes.

The CONTEXT snapshot buffer stays a NativeMemory allocation rather than
stackalloc: CONTEXT requires 16-byte alignment, now documented at the
call site. Native call sequences are unchanged.

* [Host] Address Copilot review: harden failure paths, honor injected platform

- Free the handler thunk page when the RX protection downgrade fails
  (the leak predates this branch, but the failure path is boot-fatal so
  releasing the page is unobservable).
- TraceThreadMode and the static diagnostics helpers now resolve host
  primitives through the backend bound to the current thread, falling
  back to HostPlatform.Current only when no run is active (identical on
  supported configs, honors injection everywhere a backend exists).
- HostPlatform.Create additionally requires an x64 process so native
  Windows ARM64 fails with the promised PlatformNotSupportedException
  instead of emitting x86-64 stubs into an ARM64 process.
2026-07-15 03:15:36 +03:00
anesr5 2a9a261913 loader: support ps5 SELF and validate ELF signatures (#157)
Co-authored-by: anes <anesrachedi@outlook.fr>
2026-07-15 01:31:16 +03:00
Mike Saito caf859cc52 Fix guest shutdown when VideoOut window is closed (#184)
Propagate Silk window close to runtime teardown so audio and CPU workers stop instead of continuing after the presentation window is dismissed.
2026-07-15 00:52:01 +03:00
Spooks d8397b022e Performance Improvements and Optimization Tweaks (#156)
* Improve Gen5 rendering performance and compatibility

* Pin .NET SDK for locked restore

---------

Co-authored-by: Spooks4576 <Spooks4576@users.noreply.github.com>
2026-07-14 20:22:52 +03:00
Mike Saito e80f96ecf5 Align SysAbi export names with Aerolib NID catalog (#137) 2026-07-14 17:10:57 +03:00
Dafenx 1d33ef90fc Harden param.json metadata parsing (#134)
Co-authored-by: Dafenx <196083014+Dafenxz0@users.noreply.github.com>
2026-07-14 17:10:32 +03:00
Spooks 787d3a1efb Fix Gen5 boot and restore stable AGC rendering (#139)
Co-authored-by: Spooks4576 <Spooks4576@users.noreply.github.com>
2026-07-14 16:13:42 +03:00
Kushida 884584da67 fix: restore WaitSema loop guard boundary (#133) 2026-07-14 15:07:25 +03:00
Spooks d43edc865a Agent/fix gen5 thread agc compat (#130)
* Fix Gen5 thread and AGC compatibility

* Trim compatibility comments

* Report selected Vulkan GPU

* Clean up CPU title label

* Improve emulator frame pacing and performance

* Regenerate package locks with pinned SDK

---------

Co-authored-by: Spooks4576 <Spooks4576@users.noreply.github.com>
2026-07-14 14:41:42 +03:00
Berk cf6964710a [emulator] Improve emulator performance by optimizing memory access and reducing unnecessary overhead in kernel and CPU execution paths (#131) 2026-07-14 14:28:44 +03:00
Spooks 63b440efcd [HLE] Fix guest-thread sync and boot for Unreal Engine titles (#102)
* [HLE] Fix guest-thread sync and boot for Unreal Engine titles

Silent Hill: The Short Message (and other UE titles) now boot the full
engine thread graph instead of hanging early. Four related fixes:

- pthread cond/mutex semantics: retain a signal raised with no waiter as
  pending, and key block/wake on the state's identity rather than a
  resolved address that could differ between lock and unlock. This ends
  the ~1.5M-call cond_wait busy-spin.

- Warm HLE type initializers and force-JIT their methods on a host thread
  at Freeze(). A .cctor or first-time JIT running on a guest thread's
  hijacked stack fail-fasts the CLR as "Invalid Program: attempted to
  call a UnmanagedCallersOnly method from managed code".

- Guest thread scheduling: pump after a wake so a readied thread actually
  runs, add a dispatcher thread for when every guest thread is parked,
  and make the pump-depth guard an atomic CAS.

- Route mutex/rwlock lock/unlock off the non-blocking leaf-import fast
  path so a contended lock can deschedule its guest thread.

Ported from the unreal-boot-fixes branch.

* [HLE] Keep mutex/rwlock unlock on the leaf-import fast path

The previous change routed all mutex/rwlock lock and unlock NIDs off the
leaf fast path so a contended lock could deschedule its guest thread. But
unlock never blocks, and taking it off the fast path made it slow enough
that Demon's Souls' job workers livelocked in a guest spinlock (millions
of mutex_unlock calls, no import progress, main thread stuck in
sceKernelWaitEventFlag).

Only *lock* needs to leave the leaf path. Restore the four unlock NIDs
(mutex + rwlock) so guest spinlocks stay cheap, while lock/rd/wrlock
remain off it for the blocking case Silent Hill needs.

* [HLE] Gate pthread_mutex_lock guest-thread blocking (fixes Demon's Souls)

Re-enabling cooperative deschedule on a contended pthread_mutex_lock
regressed Demon's Souls: its job workers run on libSceFiber, and blocking
a guest thread mid-fiber left sceFiberSwitch returning ESRCH followed by
a null fiber-context deref (0xC0000005). Bisect confirmed the pthread
change as the cause; the game reaches the same point as before it once
the block is skipped.

Gate the block behind SHARPEMU_MUTEX_LOCK_BLOCKING (off by default) so
contended locks fall through to the synchronous host-thread wait. The
rest of the pthread fixes (cond_wait pending signals, identity wake keys)
are unaffected.
2026-07-13 18:59:38 +03:00
Digote 46b729c5b4 [CPU] Preserve guest return value across TLS lookup (#104)
Co-authored-by: diego <diego@DIGOTE-PC>
2026-07-13 17:22:53 +03:00
Digote 3c2134474d [HLE] Avoid duplicate export registration (#100)
Co-authored-by: diego <diego@DIGOTE-PC>
2026-07-13 17:18:21 +03:00
Mike Saito 3d2c30b151 fix(core): lazy dlsym stub materialization, COW snapshots and deferred bootstrap logging (#94)
* fix(core): implement 4-tier lazy dlsym stub materialization and argument normalization

Enforce transactional and thread-safe resolution for standalone ELF bootstrapper pipelines. - Implement a 4-tier additive fallback cascade (T0: runtime symbols, T1: import entries scan, T2: Aerolib mapping, T3: runtime slack-pool lazy stub allocation at 0x7000_0000_0000). - Fix UnmanagedCallersOnly CLR runtime crashes on second bootstrap by adding NormalizeKernelDynlibDlsymArguments to detect and swap mirrored (symbol, handle) register inputs via rigorous pointer bounds verification. - Protect failure paths via CompleteKernelDynlibDlsymFailure, cleanly zero-filling target outputAddress buffers and returns Rax = -1 with zero managed logging execution in hot native paths.

* fix(core): harden lazy-stub diagnostics with COW snapshots and deferred bootstrap logging

Follow-up to lazy stub pool copy-on-write publishing in TryGetOrCreateLazyImportStub. - Snapshot _importEntries in ProbeReturnRip before near-call and PLT import lookup loops to prevent torn iteration during concurrent array replacement. - Defer SHARPEMU_LOG_BOOTSTRAP output: hot path records raw register slots in a ring buffer under lock; TryReadAsciiZ and Console.Error run only after import handler completion via DrainDeferredBootstrapTraces. - Normalize bootstrap dynlib register order at DispatchImport gateway entry before any logging or trace reads, so swapped RDI/RSI on Import#2 cannot fault the native gateway when bootstrap tracing is enabled. - Resolve lazy stub pool bounds from the full SelfLoader-mapped import region via VirtualQuery instead of a hardcoded 4 KiB cap. - Use ConcurrentDictionary for runtime symbol registration during concurrent dlsym. - Emit distinct [LOADER][WARN] reasons when import stub region resolution fails versus lazy stub pool exhaustion.
2026-07-13 12:25:37 +03:00
Berk 61d28e9e08 [CPU] optimize strcasecmp for hot path (#95) 2026-07-13 12:19:57 +03:00
Digote 4bd42795c7 [logging] Migrate HLE diagnostics to SharpEmuLog (#80)
Signed-off-by: Digote <45742711+Digote@users.noreply.github.com>
2026-07-13 12:11:27 +03:00
kuba 6e2878f2ff Add commit hash to video window title (#93) 2026-07-13 11:21:41 +03:00
Foued Attar cdee77521e Fix UnmanagedCallersOnly boot crash & Core Engine Improvements (CPU, HLE, AGC) (#81)
* [agc] WAIT_REG_MEM suspend/resume, draw packet fixes, new HLE exports, debug cleanup

Rebased onto upstream 79a7437 (par274/sharpemu, rewritten history).

- GpuWaitRegistry: DCBs suspended on unsatisfied WAIT_REG_MEM are re-polled
  against guest memory on every submit; fixed 64-bit and standard packet parse
  offsets, apply the mask, treat PM4 compare function 0 as "always".
- TryReadSubmittedDrawCount: accept the 5-dword ItDrawIndex2 form emitted by
  DcbDrawIndex (count at +4); menu draws were silently discarded before.
- sceAgcDriverSubmitMultiDcbs: reversed ABI (rdi=address array, rsi=dword
  sizes, rdx=count).
- VideoOut: vblank events, sceVideoOutGetFlipStatus, buffers registered via
  sceVideoOutRegisterBuffers are valid flip targets.
- New HLE: libc stdio (fopen/fread/fseek/ftell/fclose/fgets), Dinkumware
  _Getpctype ctype table, NpTrophy2 stubs, AMPR PAK sequential-read tracker,
  MsgDialog lifecycle, NGS2 alt NIDs + dummy vtable for handle objects,
  guarded memset intrinsic, abort()/strcasecmp null-arg recovery.
- Removed investigation-only code (INT3 breakpoints, qfont/mcpp dumps,
  error-candidate printf traces, unconditional debug logs).

First rendered frame: Quake (PPSA01880) presents a 1920x1080 guest frame.

* Implemented a guarded native intrinsic (rep movsb) in DirectExecutionBackend to bypass HLE dispatch overhead, while preserving memory safety checks.

* [hle] clock_gettime clock ids, AudioOut2 canary fix, NID rebinds, new offline stubs

- clock_gettime (lLMT9vJAck0): support CLOCK_SECOND and the *_PRECISE/*_FAST
  variants instead of returning EINVAL, which games treated as fatal and
  retried in a tight loop.
- AudioOut2: context param writes shrunk to the guest-observed layout (the
  old 0x80-byte reset smashed the stack canary at +0x60 and killed audio
  init); ContextQueryMemory writes the single u64 the caller expects.
- NGS2: dropped wrong alt-NID aliases (they hash to sceImeUpdate,
  sceMouseRead, sceSystemGestureUpdateAllTouchRecognizer - now bound in
  their real libraries); added sceNgs2PanInit; fixed VoiceGetState NIDs.
- New verified stubs: sceUltInitialize, sceNpUniversalDataSystemDestroyHandle,
  sceNpGetOnlineId, sceNpGetNpReachabilityState, sceImeKeyboardOpen,
  sceImeKeyboardGetResourceId, sceMouseOpen, sceKernelAprGetFileSize.
- Import gateway: unwind guest workers at dispatch during backend teardown;
  env-gated SHARPEMU_LOG_THREAD_MODE tracing.

* [cpu] Isolate guest execution on native worker threads

Guest entry stubs no longer run above CLR-managed frames: each run is handed
to a pooled raw OS thread whose loop is emitted native code. While guest code
executes there is not a single managed frame on the thread and it stays in
preemptive GC mode, so the GC never walks a frame chain interleaved with
guest stubs that carry no CLR unwind info (the ReversePInvokeBadTransition /
UnmanagedCallersOnly FailFast class of crashes on pumped guest threads).

- NativeGuestExecutor: CreateThread + emitted run loop (WaitForSingleObject,
  UnmanagedCallersOnly prologue/epilogue, entry stub call, SetEvent). The
  prologue rebinds guest TLS, the host-RSP slot, thread affinity and the
  Active* ambient per run, so workers carry no guest identity and pool
  freely; the orchestrating managed thread parks in a preemptive wait.
- All three entry sites route through RunGuestEntryStub: guest thread
  entries, blocked-continuation resumes, and the main ExecuteEntry.
- Teardown stops workers before any executable stub or TLS index they
  reference is freed; a worker that will not stop leaks its loop instead of
  freeing running code.
- Kill switch: SHARPEMU_DISABLE_NATIVE_GUEST_WORKERS=1 restores the inline
  calli path.
2026-07-12 18:04:12 +03:00
Foued Attar e1cf5b13ef [AGC] Quake rendering progress: WAIT_REG_MEM, draw fixes, VideoOut, and HLE improvements (#68)
* [agc] WAIT_REG_MEM suspend/resume, draw packet fixes, new HLE exports, debug cleanup

Rebased onto upstream 79a7437 (par274/sharpemu, rewritten history).

- GpuWaitRegistry: DCBs suspended on unsatisfied WAIT_REG_MEM are re-polled
  against guest memory on every submit; fixed 64-bit and standard packet parse
  offsets, apply the mask, treat PM4 compare function 0 as "always".
- TryReadSubmittedDrawCount: accept the 5-dword ItDrawIndex2 form emitted by
  DcbDrawIndex (count at +4); menu draws were silently discarded before.
- sceAgcDriverSubmitMultiDcbs: reversed ABI (rdi=address array, rsi=dword
  sizes, rdx=count).
- VideoOut: vblank events, sceVideoOutGetFlipStatus, buffers registered via
  sceVideoOutRegisterBuffers are valid flip targets.
- New HLE: libc stdio (fopen/fread/fseek/ftell/fclose/fgets), Dinkumware
  _Getpctype ctype table, NpTrophy2 stubs, AMPR PAK sequential-read tracker,
  MsgDialog lifecycle, NGS2 alt NIDs + dummy vtable for handle objects,
  guarded memset intrinsic, abort()/strcasecmp null-arg recovery.
- Removed investigation-only code (INT3 breakpoints, qfont/mcpp dumps,
  error-candidate printf traces, unconditional debug logs).

First rendered frame: Quake (PPSA01880) presents a 1920x1080 guest frame.

* Implemented a guarded native intrinsic (rep movsb) in DirectExecutionBackend to bypass HLE dispatch overhead, while preserving memory safety checks.
2026-07-12 00:22:48 +03:00
kostyaff 8a40251a1c [logging] Migrate SelfLoader and DirectExecutionBackend.Diagnostics to SharpEmuLog (#65)
Phase 2 of Console.Error.WriteLine → structured logging migration.

SelfLoader.cs (30 sites):
- Category: "Loader"
- TLS load_start/load_done, Segment info, DTPMOD64 patching → Debug
- ELF alignment mismatch, invalid symbol value skips, CRITICAL invalid patch → Warning
- Runtime symbol index populated, Initializers discovered → Info
- [FOCUS][SCAN/SKIP] relocation trace, [RELOC] target trace → Debug
- TryLoadTableBytes diagnostics → Debug (FAILED → Warning)
- ResolveMappedAddressOrFallback trace → Debug

DirectExecutionBackend.Diagnostics.cs (17 sites):
- Category: "Native" (Log field in main DirectExecutionBackend.cs partial)
- DumpRecentImportTrace → Info
- Suspicious unresolved pointer hits/cap → Warning
- ProbeReturnRip return-rip bytes/slots/PLT trace → Debug

Level mapping: [TRACE]/[FOCUS]/[RELOC]/[TEST] → Debug; [INFO] → Info;
[WARNING]/WARNING/CRITICAL/Skipping/FAILED → Warning.

[LOADER] prefix dropped — category is in LogEntry.

Console.WriteLine (stdout, ~25 sites in SelfLoader.cs) intentionally
left untouched — only Console.Error.WriteLine was in scope.

Build: 0 errors, 0 warnings.

Co-authored-by: Hermes Atlas <hermesatlas@example.com>
2026-07-11 22:17:40 +03:00
PandaCatz f43f7cde9c [cpu] Implement SysV variadic float ABI (xmm0-7 capture, float returns, printf %f) (#59)
* [cpu] Implement SysV variadic float ABI (xmm0-7 capture, float returns, printf %f)

The import trampoline spilled only xmm0 and never reloaded a return xmm0. The
guest uses the System V AMD64 ABI: variadic float args pass in xmm0..xmm7 and
float/double returns come back in xmm0. As a result variadic float args past
the first were unavailable to HLE handlers, float returns never reached the
guest, and direct printf read %f/%e/%g from GP registers instead of XMM,
printing garbage and desynchronizing every following argument.

- Trampoline: spill xmm0..xmm7 into a 0x80-byte save area below the GP argpack
  (r12 stays at the argpack base) and reload the return xmm0 in the epilogue.
- Gateway: read xmm0..7 from the save area into CpuContext and write the
  handler's xmm0 back. XMM is caller-saved in SysV, so restoring xmm0 on return
  is safe for non-float imports too.
- RegisterPrintfArgumentSource: read float args from xmm0..7 with independent
  GP/FP counters and a shared stack-overflow cursor.

Every emitted byte was decoded; a unit test confirms float args read xmm0..7
(not GP) and interleaved "%d %f %d %f" stays synchronized. Build 0/0.

* [cpu] Document the scalar-only leaf-import constraint at its registration site

- IsLeafImport: spell out the no-XMM-args / no-XMM-return invariant the fast
  path relies on and what breaks if it is violated; record the 2026-07-11 audit.
- Name every previously uncommented NID in the leaf list (mutex lock/unlock,
  usleep, the Ampr/Apr command-buffer block, the unknown AGC packet NID).
- IsNoBlockLeafImport: document that it is a sub-filter of IsLeafImport and
  that its five extra entries currently take the full gateway path; fix the
  mislabeled K-jXhbt2gn4 comment (pthread_mutex_trylock, not
  scePthreadMutexTrylock, which is upoVrzMHFeE).
- Point the DispatchImport call-site note at the audited list.

Comment-only change: the comment-stripped diff is empty and the solution
builds with 0 warnings / 0 errors.
2026-07-11 17:41:25 +03:00
j92580498-max ef74680167 More logger improvements (#58)
* fix

* fix

---------

Co-authored-by: j92580498-max <j92580498-max@users.noreply.github.com>
2026-07-11 17:34:49 +03:00
kostyaff d9d1aeaef9 [logging] Migrate CpuDispatcher, SharpEmuRuntime, PhysicalVirtualMemory to SharpEmuLog (#51)
Replaces 34 Console.Error.WriteLine call sites across 3 Core files
with structured SharpEmuLog calls (Debug/Info/Warning/Error/Critical).

CpuDispatcher.cs (10 sites):
- DispatchEntry/DispatchModuleInitializer START and entry-point logs -> Debug
- FATAL EXCEPTION catch blocks -> Critical (with exception object)
- Native backend FAILED -> Error

SharpEmuRuntime.cs (23 sites):
- Loading/Entry/Dispatching/DispatchEntry returned -> Info
- Module load/registered/preload summary -> Info
- Initializer dispatch failed/module start failed -> Error
- Imported data unresolved -> Warning, write-failed -> Error
- Import stub conflict -> Warning
- Trace-level rebind logs -> Debug

PhysicalVirtualMemory.cs (1 site):
- TraceVmem helper -> Log.Debug (SHARPEMU_LOG_VMEM env gate preserved)

Build: 0 errors, 0 warnings

Co-authored-by: Hermes Atlas <hermesatlas@example.com>
2026-07-11 11:54:15 +03:00
kostyaff c0fd6a80e8 Astro Bot shader type 4, pthread_cond_timedwait, and HLE/memory/cpu bug fixes (#40)
* [agc] Add shader type 4 (GS) and register defaults v13 support

Astro Bot (#11) crashes on boot due to two missing GPU features:

1. Shader type 4 (Geometry Shader) — SPI_SHADER_PGM_LO/HI register
   offsets 0x8A/0x8B were missing. Added constants and switch cases
   for shader type 4 in GetExpectedSpiShaderPgmLo/Hi. Also added
   type 4 to IsEsGeometryShaderType (2 or 4 or 6).

2. Register defaults version 13 — was not recognized as supported.
   Added RegisterDefaultsVersion13 constant and included it in
   IsSupportedRegisterDefaultsVersion.

* [kernel] Add POSIX pthread_cond_timedwait export

SILENT HILL (#4) and Poppy Playtime (#3) crash on boot due to
missing POSIX pthread_cond_timedwait (NID 27bAgiJmOh0).

The Sony wrapper scePthreadCondTimedwait (NID BmMjYxmew1w) was
already implemented, but the raw POSIX symbol was not exported.
Added [SysAbiExport] for pthread_cond_timedwait delegating to
existing PthreadCondWaitCore with timed: true.

* [memory] Fix FlushInstructionCache null process handle

PhysicalVirtualMemory.cs called FlushInstructionCache with null as the
process handle in two places (SetProtection and TryWriteExclusive).
On Windows, a null handle does not reliably resolve to the current
process — the correct call is GetCurrentProcess() (pseudo-handle -1).

Also corrected the P/Invoke signature:
- Changed return type from void to bool with [return: MarshalAs(Bool)]
- Added SetLastError = true
- Added GetCurrentProcess() P/Invoke import

This matches the pattern already used in DirectExecutionBackend.cs
which correctly passes GetCurrentProcess() to all FlushInstructionCache
calls.

* [hle] Distinguish NOT_FOUND from NOT_IMPLEMENTED and log duplicate NIDs

Three diagnostic improvements to the HLE dispatch path:

1. ModuleManager.RegisterFromAssembly — duplicate NID registration was
   silently skipped (dispatchTable first-wins, exportTable last-wins,
   causing metadata divergence). Now logs a warning with the NID and
   export name so conflicts are visible.

2. ModuleManager.TryDispatch — generation mismatch returned
   ORBIS_GEN2_ERROR_NOT_FOUND, conflating 'function does not exist'
   with 'function exists but not for this generation'. Now returns
   ORBIS_GEN2_ERROR_NOT_IMPLEMENTED for generation mismatch, matching
   the existing convention in CpuDispatcher. Also adds debug logging
   for both NOT_FOUND and NOT_IMPLEMENTED paths.

3. DirectExecutionBackend.Imports.cs — the import dispatch else-branch
   (the actual hot path that bypasses ModuleManager.TyDispatch via
   cached export) had the same conflation. Split into:
   - else if (export exists but generation mismatch) → NOT_IMPLEMENTED
   - else (no export at all) → NOT_FOUND
   This makes runtime diagnostics correctly distinguish missing exports
   from generation-unsupported exports.

* [cpu] Check VirtualProtect return values in all stub creation paths

9 VirtualProtect calls in DirectExecutionBackend.cs had unchecked
return values. If VirtualProtect silently fails, memory protection
remains incorrect — stubs allocated with PAGE_EXECUTE_READWRITE (0x40)
never get downgraded to PAGE_EXECUTE_READ (0x20), or guest thread
entry stubs never get upgraded to writable. This causes access
violations on next execution or silent data corruption.

Fixed all 9 sites with proper error handling:
- 6 stub creation methods (return 0 on failure + log error)
- 2 guest thread entry methods (set reason + return Exception)
- 1 guest entry method (set LastError + return MEMORY_FAULT)

Stub creation sites fixed:
- CreateImportDispatchStub (line ~1683)
- EnsureTlsHandler (void, log + return)
- CreateUnresolvedReturnStub (return 0)
- CreateGuestReturnStub (return 0)
- CreateExceptionHandlerTrampoline (return 0)
- CreateTlsStoreHelperStub (return 0)

Guest thread entry sites fixed:
- StartGuestThreadNativeCall (return Exception)
- StartGuestContinuationNativeCall (return Exception)
- RunGuestEntryPoint (return MEMORY_FAULT)

* [kernel] Remove unused duplicate _nextFileDescriptor field

KernelExports.cs declared _nextFileDescriptor but never used it.
The actual field used for file descriptor allocation lives in
KernelMemoryCompatExports.cs (lines 1314, 1337). This was a dead
duplicate causing CS0414 warning.

Build is now 0 errors, 0 warnings.

---------

Co-authored-by: Hermes Atlas <hermesatlas@example.com>
2026-07-10 21:48:50 +03:00