Compare commits

..

196 Commits

Author SHA1 Message Date
ParantezTech c6430bc51d [readme] update DeS screenshot 2026-07-23 03:05:41 +03:00
ParantezTech c66d7a1e05 [Bink2] rework bridge to use FFmpeg's native Bink2 decoder instead of a C bridge 2026-07-23 03:01:10 +03:00
Mariano Zambelli 559b7f0a84 feat(voice): add QoS stubs (GetStatus, Terminate, SetMode) (#541)
* feat(voice): add QoS stubs (GetStatus, Terminate, SetMode)

Titles call these functions during voice/multiplayer setup to check
network availability and configure modes. Unresolved imports caused
WARN floods in the loader logs. Reporting initialized + disconnected
lets callers take their normal offline path.

* fix(voice): return success (0) from sceVoiceQoSGetStatus instead of disconnected state
2026-07-23 01:48:55 +03:00
MarcelMediaDev 2272b9b576 fix(ajm): silence BatchJobDecode/Start/Wait/Cancel hot-path stubs (#547)
Unresolved batch NIDs flooded Import WARNs on Bink/AJM. Claim input
consumed with silence produced; this is not a real codec.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-23 01:44:07 +03:00
MarcelMediaDev 8dd3172c0f fix(systemservice): stub notice-screen skip flag setters (#549)
Settings probes Set/DisableNoticeScreenSkipFlagAutoSet; unresolved
NOT_FOUND can stall the SaveModTime/Load path.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-23 01:41:08 +03:00
Berk e7ea186ea8 Enhance contribution guidelines with PR expectations
Added expectations for pull requests regarding observable behavior and testing requirements. Clarified guidelines for AI-assisted contributions.
2026-07-23 01:40:00 +03:00
MarcelMediaDev 74a519875b fix(agc): add missing Cb/Dcb GetSize stubs for packet sizing probes (#535)
Unresolved GetSize NIDs returned NOT_FOUND during RenderThread startup,
leaving null packet pointers and an immediate write AV. Return fixed
packet byte sizes in rax only — no guest memory writes.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-23 00:18:47 +03:00
Andrey Modnov 4682e64e81 [CLI] Simplify mitigated child arguments (#529) 2026-07-23 00:18:11 +03:00
Kurt Himebauch 912883de05 fix(cmake): invalidate stale FFmpeg library cache (#543) 2026-07-22 23:56:37 +03:00
Berk f704586a8d [VideoOut] Add Bink2 support via FFMPEG bridge (#527)
* [VideoOut] Add Bink2 support via FFMPEG bridge

* [CMake] update commit

* [CMake] update commit
2026-07-22 21:41:41 +03:00
jute-ado d3600c9255 fix(ajm): accept Gen5 codec types (#526) 2026-07-22 18:24:06 +03:00
jute-ado 5f97031df5 shader: allow larger bounded Gen5 programs (#514) 2026-07-22 14:46:15 +03:00
h4sht 2a4da8c0a9 [Kernel/Semaphore] Close race between sceKernelWaitSema and sceKernelSignalSema (#504)
When sceKernelWaitSema finds the count insufficient it increments
WaitingThreads, releases the semaphore gate, and calls
RequestCurrentThreadBlock to set the thread-static block flags. A
signal arriving before the scheduler registers the block metadata
is missed by WakeBlockedThreads — the waiter has not been
registered yet and the signal's wake iteration skips it.

The scheduler's exit handler already re-checks TryWake() after
setting the thread to Blocked, but that requires the thread to
fully exit to the scheduler and back. Instead, re-check the
semaphore count under the gate immediately after the block request:
if the count is now sufficient, consume the tokens, cancel the
pending block via TryConsumeCurrentThreadBlock, and return without
ever yielding to the scheduler.

Co-authored-by: tru3 <tru3@tru3.com>
2026-07-22 14:34:19 +03:00
h4sht 4c37e64c66 [NpWebApi2] Add sceNpWebApi2PushEventCreateFilter stub (#503)
Add sceNpWebApi2PushEventCreateFilter (NID: MsaFhR+lPE4) to the
libSceNpWebApi2 module. This function is called by Unity games
during initialization and was unresolved, causing an import warning
and returning ORBIS_GEN2_ERROR_NOT_FOUND.

The stub validates the library context and returns an incrementing
filter handle, following the same pattern as the existing
sceNpWebApi2PushEventCreateHandle.

NID sourced via:
  python scripts/aerolib_catalog.py lookup MsaFhR+lPE4

Co-authored-by: tru3 <tru3@tru3.com>
2026-07-22 14:33:43 +03:00
kostyaff fc9e3ff393 fix: roll back earlier host allocations on later gap failure in TryBackFixedRange (#472) (#474)
When a fixed mapping spans multiple free runs and a later gap cannot be
backed, any earlier host allocations were leaked. Stage all allocations
during the walk and insert MemoryRegions only after every gap has been
backed successfully. On any failure, free all staged allocations.

Fixes #472

🤖 Generated with Hermes Agent
2026-07-22 14:28:25 +03:00
samto6 eb47d753f6 [Ampr] Implement the FW 4.00 write-address command exports (#510) 2026-07-22 03:00:30 +03:00
h4sht 6aa78bb55b [Loader] Fall back to fixed-range backfill when main image base is occupied (#493)
When TryAllocateAtExact fails for the main image base (0x800000000
for PS5, 0x400000 for PS4), the loader previously threw a fatal
InvalidOperationException with no recovery path. This happens when
the host OS has already claimed part of that address range — common
under Rosetta 2, with aggressive ASLR, or when another process maps
into the guest address space.

Instead of failing immediately, attempt TryBackFixedRange which
backs the range page by page, claiming any free gaps. If the
backfill also fails, Clear() rolls back partial allocations and
the exception now includes platform-specific recovery advice.

This prevents the most common emulator startup crash on affected
hosts.

Co-authored-by: tru3 <tru3@tru3.com>
2026-07-21 18:18:58 +03:00
h4sht 9be6f85ef0 [Font] Implement sceFontGetVerticalLayout (#492)
Add sceFontGetVerticalLayout (NID: 3BrWWFU+4ts) to the Font module,
completing the vertical-text counterpart to the existing
GetHorizontalLayout. The SceFontVerticalLayout structure is three
floats (baseline, lineAdvance, decorationExtent) interpreted for
vertical writing such as CJK text rendered top-to-bottom.

- Write baseline=8.0f, lineAdvance=16.0f, decorationExtent=0.0f
- Validate output pointer and return INVALID_ARGUMENT on null
- Return MEMORY_FAULT when guest writes fail

Tests:
- GetVerticalLayout_WritesExactlyThreeFloats with sentinel guard
- GetVerticalLayout_NullBuffer_ReturnsInvalidArgument

NID sourced via: python scripts/aerolib_catalog.py lookup sceFontGetVerticalLayout

Co-authored-by: tru3 <tru3@tru3.com>
2026-07-21 18:18:16 +03:00
Kurt Himebauch 4c8c67a3dd fix: Add ASTRO BOT compatibility stubs (#481)
* Add ASTRO BOT compatibility stubs

* Fix ASTRO BOT compatibility stubs
2026-07-21 18:17:40 +03:00
999sian ada67a1924 cpu: recover SSE4a EXTRQ/INSERTQ faults on Linux (#482)
The fault-time SSE4a fallback was Windows-only because the POSIX signal
bridge never carried XMM state: the CONTEXT scratch buffer only held the
17 general-purpose registers, so emulating EXTRQ/INSERTQ there would
have computed results from zeroed bytes and discarded the write. Bridge
the XMM registers on Linux by copying them between the mcontext's
FXSAVE image (kernel sigcontext ABI, libc-independent) and the CONTEXT
FltSave slots on capture and write-back, and gate the recovery on that
bridge instead of on Windows. Darwin still declines: its XMM area
remains unbridged.

With this, guest EXTRQ/INSERTQ on Linux hosts without SSE4a (any Intel
CPU) resumes with correct register state instead of dying on an
unrecovered SIGILL (#328).
2026-07-21 14:18:21 +03:00
Slick Daddy 2379e8988c [Loader] Collect stub-eligible NIDs in one pass over descriptors (#489)
BuildImportStubs filtered orderedImportNids by calling ShouldCreateImportStub
for each unique NID, and every call scanned the entire descriptor list
looking for a match. On a real module both the NID count and the descriptor
count run into the thousands, so the filter degraded to O(nids * descriptors)
ordinal string comparisons on the one-time load path.

Replace the per-NID rescan with a single pass over the descriptors that
builds a HashSet of eligible NIDs, then filter orderedImportNids with O(1)
membership. Eligibility is unchanged: a NID qualifies when any of its
descriptors is non-weak, or is weak but resolvable via the module manager.

ShouldCreateImportStub is retained (still used by the DEBUG self-checks), and
a self-check now asserts the set-based collector agrees with the per-NID rule.

Co-authored-by: slick-daddy <slick-daddy@users.noreply.github.com>
2026-07-21 12:59:30 +03:00
Slick Daddy 105c58b380 [Tests] Isolate Gen5 scalar fallback test from parallel static mutation (#488)
ScalarLoadReadsTrackedFallbackMemory swaps the process-global static
Gen5ShaderScalarEvaluator.FallbackMemoryReader under a lock private to the
test class. The SharpEmu.Libs [ModuleInitializer] (AgcShaderCompilerHooks)
assigns the same static to TryReadShaderGuestMemory the first time any Libs
type is touched, and it does not take that lock. Under xUnit's default
cross-class parallelism a concurrent Libs test could fire the initializer
mid-test, clobbering the swapped-in reader — observed on CI (linux-x64) as
the fallback returning all zeros: Expected [1181044592, 4, 1319632096, 4],
Actual [0, 0, 0, 0].

Put the test in a DisableParallelization collection, matching the existing
convention for shared-mutable-static tests (KernelMemoryCompatState,
AjmState, AvPlayerPathState). The collection runs alone in the non-parallel
phase, so no other test can mutate the static while this one holds it.

Co-authored-by: slick-daddy <slick-daddy@users.noreply.github.com>
2026-07-21 12:59:03 +03:00
Slick Daddy da35f0db47 [Audio] Hoist volume clamp out of the per-sample PCM loop (#487)
Co-authored-by: slick-daddy <slick-daddy@users.noreply.github.com>
2026-07-21 12:58:35 +03:00
Slick Daddy 1f3963c543 [Gpu] Factor the exact-XOR swizzle equation in the texture detiler (#483)
TryDetile's exact-XOR fast path (PS5 swizzle modes 5/9/24/27) ran the
full AddrLib address equation per element: a 16-bit interleave with 32
PopCount calls for every pixel of textures that are millions of elements.

Each output bit is parity(x & XMask) XOR parity(y & YMask), and parity
distributes over XOR, so the offset factors into independent xTerm(x) ^
yTerm(y) fields. Precompute the per-column X term once and hoist the Y
term per row, collapsing the inner loop to one array load and one XOR.

Add GnmTilingDetileTests, which lays out a tiled buffer from an
independent re-derivation of the mode-27 equation and asserts TryDetile
reconstructs it byte-for-byte.

Co-authored-by: slick-daddy <slick-daddy@users.noreply.github.com>
2026-07-21 12:57:39 +03:00
iExplosiveRage 4bb1af93d7 SaveData: avoid invalid DeS transaction resource pointer (#480)
Demon's Souls treats the small transaction-resource handle as a guest pointer during the fresh-save path. Return a null resource for the observed call shape to prevent the repeatable access violation at address 0x9.

Co-authored-by: RedDv <RedDv@DESKTOP-EVNB4S8>
2026-07-21 02:22:51 +03:00
Nicola Pomarico 0ae785c617 [VideoPresenter] Accept padded row pitch in guest image uploads (#475)
The guest can hand initial texture data whose rows are padded out to a
hardware alignment wider than the image width, so the total byte count
exceeds the tightly packed width*height*bpp we compute. The upload path
rejected any byte count that did not match exactly, silently dropping
these uploads and leaving the texture blank.

Recover the real source row length when the byte count is consistent
with a common padding alignment (8/16/32/64/128/256 texels) and pass it
through as BufferRowLength on the copy, instead of always hardcoding 0.
Uploads that do not match a recognised padded layout are still rejected
as before.

Verified against Dead Cells (PPSA15552): a loading-transition texture
upload that previously wedged the title now uploads correctly and the
game proceeds past the load screen, running stably past 1M draw calls
with no stalls. Dreaming Sarah (tightly packed path) still renders
normally, confirming no regression to the non-padded case.
2026-07-21 01:01:28 +03:00
Slick Daddy e01092aa38 Kernel FS: close guest→host sandbox escapes in the path resolver (#478)
* Kernel FS: default-deny unmapped guest paths (fixes absolute-path host escape)

ResolveGuestPath returned any unrecognized guest path verbatim as the host
path. Because absolute paths ("/etc/passwd", "C:\Windows\...") are already
fully qualified, they skipped the relative-path app0 fallback and were handed
straight to FileStream/File.Delete/etc., giving a malicious game arbitrary
host-file read/write/delete outside the sandbox.

Return string.Empty (deny) on fallthrough instead. Most callers already treat
a nonexistent host path as NOT_FOUND; open/truncate/rename get an explicit
empty-path guard so a denied path can't reach FileStream and throw an
ArgumentException their catch blocks don't cover.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Kernel FS: contain built-in mounts (fixes Windows drive-letter injection)

The built-in mount branches (app0/temp0/download0/hostapp/devlog) combined
the mount-relative guest path onto the host root without re-checking
containment. NormalizeMountRelativePath clamps ./.. but splits only on
separators, so a drive-qualified token like "C:" survives as a segment and
Path.Combine then discards the mount root, yielding a raw host path such as
"C:\Windows\..." (arbitrary host read/write).

Route every built-in branch through a new CombineWithinMount helper that
re-resolves with Path.GetFullPath and verifies the result stays under the
mount root -- the same guard TryResolveRegisteredGuestMount already applied.
Denied paths return string.Empty, which callers treat as unresolved.

AprStreamingContractTests passed a raw Path.GetTempFileName() as the guest
path, relying on the now-removed absolute-path passthrough; updated it to
address the file through a registered mount.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Kernel FS: reject reparse points inside mounts (fixes symlink escape)

Lexical containment (Path.GetFullPath + StartsWith) proves the textual
path stays under the mount root but does not follow symlinks/junctions.
A malicious game dump could plant a reparse point inside app0/temp0/etc.
pointing outside it, so a contained-looking path resolved onto the host
filesystem. Walk each existing component from the mount root to the
candidate and refuse any reparse point, in both the built-in and
registered-mount resolution paths. Mirrors AvPlayer's existing defense.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Kernel FS: fail closed when path containment cannot be verified

The reparse-point and drive-letter containment guards call Path.GetFullPath
and File.GetAttributes on untrusted guest paths. Both throw on crafted
over-long or invalid-char input, and ResolveGuestPath runs outside the file
syscalls' try blocks, so such a path was a guest-triggerable crash rather
than a denial.

Wrap the GetFullPath calls in CombineWithinMount and the registered-mount
path, and widen the GetAttributes catch, to treat any access/format failure
as an escape (deny) instead of propagating. Also tighten the ".." fallback
check so a legitimate file named "..foo" is not falsely rejected, and hoist
the repeated Path.GetFullPath(mountRoot) into a local.

Adds a regression test asserting the resolver returns without throwing for
an over-long and a NUL-embedded path under a mount.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Kernel FS: assert malformed paths resolve to empty, not just no-throw

The fail-closed regression test asserted only Assert.NotNull, which a
non-nullable string return can never violate via its value (only a throw,
which aborts the test earlier anyway). Tighten to Assert.Equal(string.Empty)
so it also locks in fail-CLOSED: a regression where a malformed path resolved
to a non-empty host path would now be caught.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Kernel FS: route new AMPR batch tests through a registered mount

Merging main brought in three AprStreamingContractTests that pass raw
Path.GetTempFileName()/temp host paths as guest paths. The default-deny
resolver from this branch rejects absolute host paths, so MissingMidBatch
failed at index 0 instead of the intended index 1. Address the present
file through a registered mount (as ResolveStatAndReadFile already does)
so entries 0 and 2 resolve and the batch fails at the genuinely-missing
entry. The two all-missing tests were unaffected but share the fix's intent.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Kernel FS: make a matched-mount denial terminal; fix Unix-only test asserts

A registered mount that claims a path by prefix but denies it (failed
containment or a reparse point inside the mount) now short-circuits in
ResolveGuestPath instead of falling through to the built-in mount branches.
The fall-through let an overlapping prefix (a registered "/app0" vs the
built-in SHARPEMU_APP0_DIR branch, which resolves against a cached root)
re-resolve a denied path and turn the denial back into a resolution -- the
reparse-point escape reappeared on Linux CI through exactly this path.

Also fix two tests that asserted Windows-specific behavior unconditionally:
a "C:\..." path is not absolute on Unix (it resolves contained under the
mount there), and that case is now pinned to Windows only.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: slick-daddy <slick-daddy@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-21 00:58:34 +03:00
TarkusTK 224a36eba7 [Gpu] Stop retrying array uploads that overrun their allocation (#476)
A 2D-array texture whose Depth times the per-slice stride runs past its
real allocation fails a slice read partway through the upload loop, and
falls through to the single-slice path after already detiling the layers
it did read. That fall-through builds the texture with ArrayLayers
defaulting to 1, so the presenter caches it under a one-layer key while
the next draw looks it up with ArrayLayers = Depth. The two never match,
so the texture misses the cache and repeats the whole read-and-detile on
every draw, throwing the result away each time.

Detiling is per-texel swizzle math, so one such texture retried a few
times per frame is expensive: it measured 568-879 ms of every second in
Demon's Souls, against a 1.4 second frame.

An allocation that is too short stays too short, so remembering the
address and not retrying it costs nothing and repairs the cache key as a
side effect: with the array upload skipped, arrayUploadLayers is 1, which
is exactly what the fall-through texture reports.

Tested on Demon's Souls (PPSA01342): 0.7 fps to 3.6-4.1 fps, CPU detile
time per second from ~700 ms to 0, and arrayed textures go from missing
the cache on every draw to hitting it every time. 28 of the 29 array
uploads in that run already succeeded and are unaffected; only the one
overrunning texture now falls back to its base slice. 495 tests pass.
2026-07-20 19:22:19 +03:00
Berk ac883e44fa [VideoPresenter] Fix logical width/height calculation (#473) 2026-07-20 16:57:38 +03:00
TarkusTK 25d741b35b [Gpu] Sample 2D array textures with real layers (#471)
Texture arrays were uploaded and viewed as plain 2D images, so every
layer index in a shader resolved to slice 0. The guest-texture gate
also rejected array, 3D and cube descriptor types outright, sending
those resources to the 1x1 black fallback.

UI atlases hit this constantly, since they pack several sheets as array
slices and pick one per vertex. In Demon's Souls the settings menu
bottom bar stretched a mid-atlas crop across itself, and slider thumbs
and button prompts drew the wrong sheet.

The MIMG decoder already had Dimension and IsArray, so
IsArrayedImageBinding makes one rule out of them for the SPIR-V
translator and the Vulkan backend to share. Both have to agree or the
declared image type and the bound view type mismatch. Sample and gather
bindings with an array address now declare an arrayed image and pass
(u, v, slice). AgcExports reads every slice at the per-slice mip-chain
stride and passes the layers packed in one buffer, which uploads as a
2D array image in a single copy region.

Load and store bindings are unchanged. Arrayed bindings that resolve to
a fallback or to a single-layer guest image get a one-layer 2D array
view so the descriptor still matches the shader.

Tested on Demon's Souls (PPSA01342): the bottom bar, slider thumbs and
button prompts draw their correct sheets. 470 tests pass.
2026-07-20 15:44:06 +03:00
TarkusTK dce7c87c4d [AGC] Implement the owner-scoped resource unregister exports (#469)
sceAgcDriverUnregisterOwnerAndResources (ZLJk9r2+2Aw) and
sceAgcDriverUnregisterAllResourcesForOwner (SCoAN5fYlUM) were
unresolved. We already register owners and resources, and the guest
registers a resource owner per streaming batch, so with no way to
release one the fixed owner pool filled up: sceAgcDriverRegisterOwner
started failing and the guest logged its own "Agc registerOwner error:
0x80020003", after which it kept half-registering records. Demon's Souls
then crashed scanning that registry.

Owner-scoped teardown is straightforward because RegisteredAgcResource
already carries its owner, so both entry points share one sweep over the
resource table. UnregisterOwnerAndResources additionally drops the owner
itself and its compute queue, and reports INVALID_ARGUMENT for an owner
that was never registered. The existing single-resource
sceAgcDriverUnregisterResource (pWLG7WOpVcw) is unchanged.

Both NIDs are checked against their export names by the SHEM004
analyzer, which fails the build on a mismatch.

Tested on Demon's Souls (PPSA01342): the registerOwner error no longer
appears and the registry stays consistent across streaming batches. 470
tests pass.
2026-07-20 15:43:47 +03:00
TarkusTK 6ee445f0c2 [AGC] Read mip 0 from its GFX10 mip-chain offset (#470)
GFX10 stores a mip chain smallest-first: the mip tail packs into the
first swizzle block, the remaining mips follow in decreasing size, and
mip 0 ends up at the end of the allocation. We read the base level
straight from the descriptor address, so every mipped sampled texture
decoded as a collage of its own smaller mips - in Demon's Souls that
showed up as scrambled menu text and repeated controller icons.

GnmTiling.TryGetBaseMipPlacement ports the AddrLib chain-offset math
from Gfx10Lib::ComputeSurfaceInfoMacroTiled/MicroTiled. It returns a
byte offset to mip 0, or, when the whole chain fits inside the tail
block, the element coordinates of mip 0 within that block.
TryCreateGuestDrawTexture applies the offset to the sampled and storage
guest reads, and TryDetileTextureSource lifts a tail-resident mip 0 out
of the detiled block as a sub-rectangle.

MAX_MIP is only decoded from extended descriptors, so resources without
one, and single-level resources, keep the current behaviour.

Tested on Demon's Souls (PPSA01342): menu text and icons decode
correctly instead of showing shrunken copies of themselves. Verified
offline by dumping the raw tiled bytes and the detiled output for a
4096x4096 UI atlas and checking the art lands at the sampled
coordinates.
2026-07-20 15:27:46 +03:00
StealUrKill 9d187dec55 Prevent AvPlayer movie startup failures across supported hosts (#456)
* Prevent AvPlayer movie startup failures across supported hosts

* Prevent GR2 startup stalls during APR file checks and adaptive mutex self-locks.
2026-07-20 15:27:37 +03:00
kuba 3574a3b145 Shader: lower VOP3P V_FMA_MIX_F32/LO/HI (was dropping Unity HDR shaders) (#466)
The decoder recognises the VOP3P mix ops (0x20 V_FMA_MIX_F32, 0x21
V_FMA_MIXLO_F16, 0x22 V_FMA_MIXHI_F16) but left them opaque
(Vop3pRaw20/21/22), so at SPIR-V emission they fell through the
vector-ALU switch to the default and failed with "unsupported vector
opcode". A single unhandled instruction fails the whole compile, so any
shader using fma_mix was dropped entirely. Unity's built-in-RP /
PostProcessing v2 HDR, tone-mapping and auto-exposure shaders emit
V_FMA_MIX_F32, so those passes never translated (this is what kept
Superliminal's auto-exposure luminance chain from running).

Name the three opcodes in DecodeVop3p (like the packed v_pk_* ops) and
lower them in the SPIR-V translator. Each mix op computes a single f32
fma(a, b, c) where every source is read *independently* as either a full
f32 register/constant or one f16 half widened to f32. Per operand,
op_sel_hi selects f16-vs-f32 and op_sel picks which f16 half; the neg_hi
field is repurposed as an absolute-value modifier and neg negates,
applied abs-then-neg. This reuses the VOP3P op_sel/op_sel_hi/neg/neg_hi
bit layout with the mix-specific meaning, not the packed-math meaning.
The result is a scalar f32 for V_FMA_MIX_F32; _MIXLO/_MIXHI narrow it
back to f16 (exact round-to-nearest-even, via the existing
EmitFloatToHalf) and write it into the low/high 16 bits of vdst,
preserving the other half. The clamp modifier saturates to [0, 1]
consistently with the other VOP3P ops. Per-operand F16/F32 select and
the abs/neg modifiers follow shadPS4's GetSrcMix, the authoritative
reference for the mix semantics.

Adds Gen5FmaMixSpirvTests: assembles V_FMA_MIX_F32 (with a representative
op_sel/op_sel_hi/neg/abs) and V_FMA_MIXLO_F16 compute shaders and asserts
they translate to GPU SPIR-V without hitting the drop path and emit a
GLSL.std.450 Fma (and an FAbs for the neg_hi modifier). Both fail against
the pre-fix tree with "unsupported vector opcode Vop3pRaw20/21".
2026-07-20 14:37:40 +03:00
Job Meijer a1cbff8a9c Fix NID BHouLQzh0X0, doubled StartupStaticTlsReservation memory. Both needed to launch GTA V. (#454)
* Increased StartupStaticTlsReservation (doubled) and fixed mistake in NID BHouLQzh0X0. Now GTA V RAGE engine seems to start loading.

* fixed NID BHouLQzh0X0, this had an issue causing GTA V not to load. Also doubled StartupStaticTlsReservation.

* Removed .vscode folder and reverted global.json
2026-07-20 14:37:30 +03:00
Spooks db9b20481c Add internal render resolution scale and fix DPI Issue (#468)
* Add internal render resolution scale and fix embedded surface DPI scaling

Adds a GUI-configurable internal resolution scale (Graphics tab) that
renders offscreen color/depth targets below native guest resolution
and upscales on present, trading image quality for GPU headroom.
Storage/UAV images and sampled asset textures are left untouched, and
texture-alias/feedback-loop lookups compare against each target's
logical (unscaled) size so scaled render targets are still found
correctly when sampled back.

Also fixes the embedded game surface not filling the window: the
isolated emulator child process had no declared DPI awareness, so
Windows silently downscaled every window-geometry query it made
against the GUI-owned surface HWND by the display's DPI factor,
leaving an unfilled black margin on scaled displays.

* Remove flaky Gen5ScalarMemoryFallbackTests

* Restore Gen5ScalarMemoryFallbackTests

---------

Co-authored-by: Spooks4576 <Spooks4576@users.noreply.github.com>
2026-07-20 14:37:20 +03:00
ParantezTech 8cd46243ab Merge branch 'main' of https://github.com/sharpemu/sharpemu 2026-07-20 14:23:24 +03:00
ParantezTech 3334707f7c [CI] fix rule name 2026-07-20 14:23:04 +03:00
kuba 20eda4443c Shader: test a wave mask consumed as a per-lane predicate at the lane bit (#465)
* Shader: read a wave mask consumed as a per-lane predicate at the lane bit

A VCC/EXEC wave mask consumed as a per-lane predicate (the VCndmask
condition, a VCC/EXEC branch, or the derived _vcc/_exec bool) was tested in
single-lane emulation with a whole-word non-zero test (IsNotZero64) instead
of the current lane's bit. That is correct for comparison results (only the
lane's own bit is ever set) but wrong for bitwise-complement wave-mask idioms
(S_NOT / S_ORN2 / S_ANDN2 / S_NAND / S_NOR), which set the unused upper 63
bits: a whole-word test then reports the lane active even when its bit is
clear.

Unity's PostProcessing NaN killer does exactly this: per channel it computes
isNaN = NLT AND NGT AND NEQ (against 0), then combines the channels as
anyNaN OR NOT(v3-is-finite) via S_ORN2_B64. The complement set the upper mask
bits, so every valid pixel read as NaN and was replaced with 0, zeroing the
whole HDR scene before Bloom/Uber/tonemap. The 3D scene therefore rendered
black behind the menu while the UI survived. Extract the current lane's bit in
both single-lane and subgroup modes so IsWaveMaskActive matches the hardware.

Fixes Superliminal (PPSA06084) black 3D scene: the storage room now renders
behind the menu with natural exposure and no forced values.

(cherry picked from commit 7af6f4b6f314fe302619c0d44f4db00971c5bf24)

* test: wave-mask predicate is tested at the current lane bit

Regression test for the wave-mask lane-bit fix. Compiles a shader that
writes VCC at run time (V_CMP_EQ_F32) and asserts the emitted SPIR-V tests
the wave mask at the current lane's bit (mask & lane_bit) rather than with a
whole-word non-zero test. Fails against the previous IsNotZero64(mask) path,
which zeroed complement wave-mask idioms (S_ORN2/S_NOT, e.g. Unity's NaN
killer) across every lane.
2026-07-20 14:18:20 +03:00
Slick Daddy bb3318a503 kernel: return -1/errno from POSIX file syscalls on failure (#461)
* kernel: return -1/errno from POSIX open and fstat on failure

The POSIX-named open (wuCroIGjt2g) and fstat (mqQMh1zPPT8) exports routed
straight to the raw sceKernel* implementations, which report failure via
the 0x8002xxxx OrbisGen2Result sentinel in the return value. libc callers
follow the POSIX ABI and expect -1 with errno set, so they stored the
sentinel as a valid fd. Unity's IL2CPP file layer did exactly this while
probing the absent /app0/Media/il2cpp.usym: open returned NOT_FOUND
(0x80020002), the guest kept the sentinel as an fd, passed it back into
fstat, and eventually dereferenced a null pointer (vmovups xmm0,[rdi],
rdi=0) deep in a native .prx, crashing with 0xC0000005.

Wrap both entry points to translate a failed raw result into -1/errno,
mirroring the existing PosixStat/PosixLseek convention. Add a shared
PosixFailure helper (fstat maps a bad handle to EBADF; path calls default
to ENOENT) and route it through PosixStat too. Covered by two regression
tests reproducing the missing-file and misused-sentinel-fd cases.

* kernel: return -1/errno from POSIX close, read and write on failure

Same defect class as open/fstat: the POSIX-named close (bY-PO6JhzhQ),
read (AqBioC2vF3I) and write (FN4gaPmuFV8) exports forwarded the raw
sceKernel* core result, leaking the 0x8002xxxx sentinel to libc callers
that expect -1/errno on a bad fd. close in particular is on the crashing
Unity path, invoked on the sentinel the guest mistook for an fd.

Wrap all three through PosixFailure with EBADF as the fd-not-found errno.
Add regression tests for each, and correct the socket test that had
locked in the old raw-sentinel contract for a double close.

---------

Co-authored-by: slick-daddy <slick-daddy@users.noreply.github.com>
2026-07-20 14:17:08 +03:00
ParantezTech d151e151c2 [CI] fix zip inside zip 2026-07-20 10:14:55 +03:00
João Victor Amorim 472fc96a37 [AGC] Support the clamp modifier on packed f16 VOP3P ops (#460)
The VOP3P emitter rejected any packed op with the clamp bit set. Clamp
saturates each f16 output half to [0, 1] (and flushes NaN to 0, matching
RDNA), so games that emit clamped packed arithmetic fell back to a loud
emit failure.

Apply the saturation to the f32 result of each lane, before it is
narrowed back to f16. Because 0.0 and 1.0 are exact in both f32 and f16
and the clamp is monotonic, clamping in f32 and then rounding to f16
yields the same value as clamping the f16 result directly; for the fused
multiply-add the pre-narrowing value is the round-to-odd f32, which
preserves that equivalence through the final round-to-nearest-even. The
saturation uses ordered compares so a NaN result collapses to 0 without a
separate IsNan test.

Verification:
- The local exact-reference harness now also clamps: add, mul, and fma
  each compared against an f16-domain clamp reference (NaN -> 0, else
  [0, 1]) over directed boundary inputs and 24M random cases. 0
  mismatches, alongside the existing 34M unclamped fma cases.
- ShaderDump pk-f16 gains a clamped add and a clamped fma; all decode and
  emit.
- The exec program computes the pinned fma with clamp (both lanes exceed
  1.0, so each saturates to 0x3C00) and stores it at offset 28;
  GpuConformance checks it on device. All values match on an AMD Radeon
  RX 7700 XT.
2026-07-20 09:09:07 +03:00
Slick Daddy 33be88bdf9 memory: back the free pages of a partially-overlapping fixed mapping (#458)
A SCE_KERNEL_MAP_FIXED request whose window partially overlaps an
existing allocation was failing outright: AllocateAt reserves the whole
range in one all-or-nothing VirtualAlloc, which returns 0 on partial
overlap. The mapping call then returned NOT_FOUND while leaving the free
tail unmapped, so the guest faulted (0xC0000005) writing into it.

Add IGuestAddressSpace.TryBackFixedRange, which walks the range via the
host Query (VirtualQuery reports contiguous same-state runs) and fills
only the free sub-ranges, leaving already-backed pages untouched. This
matches the fixed-mapping contract on hardware. Route the fixed
reservation path through it via a new backPartialOverlap flag.

Co-authored-by: slick-daddy <slick-daddy@users.noreply.github.com>
2026-07-20 09:08:29 +03:00
kadu04t 184e24fbb6 PerGameSettings Null toggles (#453) 2026-07-20 01:30:13 +03:00
kuba 327018e80a Encode linear-float flips to sRGB at present (#448)
PS5 float VideoOut buffers (A16B16G16R16F flips) hold linear scRGB
light where 1.0 is SDR white; hardware scan-out applies the display
transfer function. vkCmdBlitImage converts numerically only, so
raw-blitting a linear-float guest frame into a UNORM swapchain crushes
dim scenes to near-black.

Blit float flip sources through a cached swapchain-sized sRGB
intermediate (the sRGB store performs the linear->sRGB encode), then
raw vkCmdCopyImage the encoded bytes into the same-compatibility-class
UNORM swapchain image. Swapchains that are already sRGB keep the
direct blit (their store encodes), and swapchain formats without an
sRGB counterpart keep today's raw blit unchanged.
2026-07-20 01:29:38 +03:00
kuba 04557fd250 Refresh CPU-rewritten guest textures by write generation (#447)
* Track guest CPU write generations

* Refresh CPU-rewritten guest textures by write generation
2026-07-20 01:29:30 +03:00
Spooks 90c72ebecf Fixes a Mutex Issue Preventing Some UE Titles From Booting (#451)
* Optimize guest import, memory, and pthread hot paths

* Fix UE adaptive mutex self-lock handling
2026-07-19 13:20:05 -06:00
Nekono 8ef5a54ee4 cpu: emulate AMD-only Zen 2 instructions in software (#449)
Handle immediate EXTRQ and INSERTQ as well as MONITORX and MWAITX when the host raises illegal-instruction faults. Add unit coverage for SSE4a bit-field semantics and preserve existing load-time patching.

Co-authored-by: zocomputer <help@zocomputer.com>
2026-07-19 21:57:42 +03:00
shadowbeat070 0c467e8c57 Add missing nids (#450)
* [Kernel] Implement clock_getres and the POSIX pthread_once alias

clock_getres (smIj7eqzZE8) was missing entirely. It reports 100ns, which
is the resolution clock_gettime here actually delivers via
DateTimeOffset.UtcNow, rather than claiming the 1ns a caller might
otherwise rely on. A null res pointer is accepted per POSIX.

pthread_once (Z4QosVuAsA0) needed no new logic: libKernel exports the
same routine under two NIDs and only scePthreadOnce (14bOACANTBo) was
registered. Shipped middleware links the plain name.

Both are imported by DOOM + DOOM II (PPSA21444): clock_getres blocked
party.prx from initialising, and pthread_once is used by libcohtml,
libPlayFabMultiplayer, party.prx and the eboot.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 848f035827)

* [Libs] Implement sceAgcGetIsTrinityMode, NpReachability and Trophy2 info

Three exports DOOM + DOOM II (PPSA21444) imports and currently receives
unresolved-stub errors for.

sceAgcGetIsTrinityMode reports the base console this backend emulates. It
returns the flag in rax and writes no guest memory: the observed rdi at
the call site sits inside the AGC state block, immediately below the
shader handles the guest stores, so writing through it would corrupt live
state if that register is stale rather than an out-pointer.

sceNpRegisterNpReachabilityStateCallback accepts the callback and never
fires it, matching the existing sceNpRegisterStateCallback handling.
Reachability transitions only occur on a live PSN connection.

sceNpTrophy2GetTrophyInfo reports NOT_FOUND rather than success.
Succeeding requires filling SceNpTrophy2Details and SceNpTrophy2Data,
whose layouts are not confirmed here, and a title trusting zeroed details
would read an empty name and grade 0 as real data. NOT_FOUND is a
documented outcome callers already handle.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 181286f621)

* [Kernel] Implement the POSIX libKernel exports titles link directly

libKernel exports many routines under both sce-prefixed and plain POSIX
NIDs, and shipped middleware links the latter. These nine are imported by
DOOM + DOOM II (PPSA21444) and had no registration at all.

Aliases onto existing implementations, identical argument order:
  mprotect (YQOfxL4QfeU), munmap (UqDGjXA5yUM), setsockopt (fFxGkxF2bVo)

New:
  getpagesize reports OrbisPageSize (16 KiB), not the host 4 KiB. An
  allocator rounding to the host value produces sub-page offsets that
  every mapping call here rejects for misalignment.

  pthread_rwlock_tryrdlock/trywrlock get a dedicated non-blocking core.
  They deliberately do not reuse TryAcquireBlockedRwlock, which
  decrements WaitingWriters -- correct only for a thread that previously
  incremented it. A fresh try never did, so reusing it would consume
  another thread's waiter count and let a queued writer be skipped.

  getsockopt reads back the three options this backend tracks (SO_NBIO,
  SO_REUSEADDR, SO_ERROR) and rejects the rest rather than returning
  success with an untouched buffer the caller would treat as real.

  send maps WouldBlock onto the existing net error path.

  inet_ntop converts AF_INET/AF_INET6 and returns the destination
  pointer per POSIX, failing rather than truncating when it will not fit.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 98c6851840)

* [Kernel] Implement the POSIX mprotect, munmap and getpagesize aliases

mprotect and munmap forward to the existing sceKernelMprotect and
sceKernelMunmap; the argument order is identical, so they are plain
aliases rather than separate implementations.

getpagesize reports OrbisPageSize (16 KiB), the granularity this backend
maps and aligns against, not the host's 4 KiB. An allocator that rounded
to the host value would produce sub-page offsets that every mapping call
here then rejects for misalignment.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* [Kernel] Implement the POSIX _nanosleep symbol

libKernel exports nanosleep under two NIDs: yS8U2TGCe1A for the plain
name and NhpspxdjEKU for the underscore-prefixed _nanosleep that libc
conventionally provides alongside it. Only the former was registered.

Both are POSIX-side symbols, so this shares NanosleepCore with posix:
true - reporting failure as -1 plus errno rather than returning an
OrbisGen2Result the way sceKernelNanosleep does.

Not exercised at runtime: no title currently on this branch imports
_nanosleep, so the choice of error convention rests on it being the
same libc routine as nanosleep, not on observed behaviour.

* [Kernel] Implement the POSIX-named pthread aliases

libKernel exports each of these routines under two NIDs: a scePthread*
name and the plain POSIX name. Only the scePthread* half was registered,
so middleware compiled against POSIX headers linked an unresolved stub.

Adds the POSIX-named export for fourteen routines, each delegating to
the existing implementation:

  pthread_setprio               pthread_attr_setschedpolicy
  pthread_getschedparam         pthread_attr_setdetachstate
  pthread_attr_getschedparam    pthread_attr_setschedparam
  pthread_attr_getstack         pthread_attr_setinheritsched
  pthread_attr_get_np           pthread_attr_setguardsize
  pthread_attr_getstacksize     pthread_attr_getguardsize
  pthread_attr_getdetachstate   pthread_rename_np

Arguments are identical in both forms, and per the convention set by
scePthreadOnce's alias the POSIX name returns the same OrbisGen2Result
rather than translating to errno.

The equivalent POSIX names for mkdir, listen, accept and recv are
deliberately not included here. Those pair with sceKernelMkdir and the
libSceNet entry points, whose error convention differs from the POSIX
one, so they need a decision about error translation rather than a
straight delegation.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 21:57:05 +03:00
Youss 9ff60abb9b [Kernel] Clamp guest path traversal at the mount root (#437)
NormalizeMountRelativePath only stripped leading separators and swapped
slashes; it never resolved "." or ".." segments. The bare-relative fallback
in ResolveGuestPath did not even call it -- it combined the guest path
against the app0 root verbatim.

A guest path containing ".." therefore escaped its mount into the host
filesystem. Unreal Engine titles hit this constantly: their base directory
is <app>/binaries/<platform>, so they address content with "../../../"
prefixes that resolve back inside /app0 on real hardware. Here those walked
out of the game folder entirely -- The Invincible (PPSA06426) opened
"/app0/.." and enumerated the host's Downloads directory, listing unrelated
user files, and never located its own content tree.

Resolve "." and ".." while walking the segments and clamp the result at the
mount root, then route the bare-relative fallback through the same
normalizer. This both closes the sandbox escape and makes the engine's
relative content paths land where the title expects.

Verified on The Invincible: no resolved host path contains ".." any more,
and the title now enumerates its own content directories (content/paks,
content/movies, content/locale/*) instead of an unrelated host folder. No
regression on Dead Cells (PPSA15552): still reaches AGC rendering and
presents frames with zero mutex errors.
2026-07-19 20:34:00 +03:00
Youss 3ebfc56d4c [PlayGo] Derive the installed chunk set from the pak files on disk (#438)
A title that ships no PlayGo sidecar was reported as a single-chunk
package. That is wrong for any package whose content is split across
chunks: the title is told everything past chunk 0 is not installed, even
though a locally dumped title has all of its data present.

The Invincible (PPSA06426) ships pakchunk0..8 and asks PlayGo which of
those are available. Receiving BAD_CHUNK_ID for chunks 1..8, it re-queried
scePlayGoGetLocus for the same chunk in a tight loop that never terminated
(observed ~1000 consecutive dispatches with identical arguments).

Discover the chunk ids from the pakchunk<N>-<platform>.pak files present
under the app0 root instead. Those N are exactly the chunks the package
has, so the answer is derived from the install rather than assumed. Chunk 0
is always included, so a title with no pak files at all keeps the previous
single-chunk behaviour, and ids outside the discovered set still return
BAD_CHUNK_ID so title-side chunk enumeration still terminates.

Verified on The Invincible: the discovered set is [0..8], matching the nine
pak files, and the GetLocus retry loop no longer occurs. No regression on
Dead Cells (PPSA15552): still reaches AGC rendering and presents frames.
The existing metadata-free contract test still passes -- its app0 fixture
has no pak files, so the discovered set stays [0].
2026-07-19 20:33:30 +03:00
Youss 73e8821d5b [Kernel] Hand off mutex ownership directly to the head waiter on unlock (#439)
pthread_mutex_unlock cleared ownership (OwnerThreadId = 0) and only woke
the head waiter, relying on that woken thread to re-acquire the lock
itself. If the wake raced or was lost, the mutex was left "free but with
a queued waiter" — a state the fast-acquire path in PthreadMutexLockCore
explicitly refuses (OwnerThreadId == 0 && Waiters.Count == 0), so every
later locker, including the game's main thread, queued behind a head that
never advanced and the whole process wedged.

Grant the mutex to the head waiter directly inside unlock (the same
TryGrantMutexWaiterLocked hand-off the thread-exit cleanup already uses),
then wake it. The mutex is therefore never observable as free-with-waiter.

Verified against The Invincible (PPSA06426): forward progress jumps from
~3.5M to ~40M dispatched imports and the repeated unlock INVALID_ARGUMENT
errors disappear. No regression on Dead Cells (PPSA15552), which still
reaches AGC rendering with zero mutex errors.
2026-07-19 20:33:21 +03:00
StealUrKill bc51cc2c4d Prevent invalid SaveData writes from damaging guest memory (#444)
Add an optional write monitor so the team can find future memory damage on each supported desktop system.
2026-07-19 20:27:17 +03:00
Nicola Pomarico d7f6e3f578 [Kernel] Implement sceKernelMapDirectMemory2 (#433)
The "2" variant of sceKernelMapDirectMemory was unimplemented, so titles
that call it (seen in Gex Trilogy) got an unresolved import that returned
an error the guest then used as a mapped address.

v2 inserts a memoryType argument ahead of v1's protection, shifting the
remaining arguments down one register and pushing alignment onto the
stack. Extract v1's body into a shared MapDirectMemoryCore and route both
exports through it; v2 reads its shifted arguments and the stack alignment
and accepts the memoryType (which only selects cache/GPU attributes this
HLE does not model per mapping, so it does not affect placement).
2026-07-19 14:35:40 +03:00
cse.aadi e56e74f960 Fix space-in-path game launching on Windows (#432) 2026-07-19 14:25:43 +03:00
kadu04t 0f224ec036 Gui Settings Null list Entries (#430) 2026-07-19 13:53:56 +03:00
kostyaff 85dc98dedc test: add Fiber exports contract tests (13 tests) (#428) 2026-07-19 13:53:33 +03:00
wearr 5d7d8e0edd [Kernel] add NID B5GmVDKwpn0 (pthread_yield) (#426) 2026-07-19 13:49:51 +03:00
wearr a60bfc9c83 [Kernel] Implement pthread semaphore exports (#424) 2026-07-19 04:18:16 +03:00
Berk 0b83b34cda chore: bump version to 0.0.2-beta.4 (#423) 2026-07-19 03:42:52 +03:00
Adam salem 09812600a0 Add libc heap trace contract tests (#409) 2026-07-19 03:25:42 +03:00
João Victor Amorim 3005babab8 [AGC] Emit v_pk_fma_f16 with exact single rounding (#420)
Completes the fused-FMA slice deferred by the VOP3P first slice (#145).
v_pk_fma_f16 previously failed emission loudly because an f32
multiply-add followed by an f16 pack rounds twice; the pinned miss is
fma(0x4100, 0x7522, 0x04EA) = 0x7A6B fused vs 0x7A6A via f32.

The f32 product of two f16 values is exact, so only the addition needs
correcting: compute sum = RN(product + addend), recover the exact
residual with Knuth 2Sum, and if the sum is inexact with an even
significand, step one ulp towards the true value. That is round-to-odd,
and rounding the f32 result to f16 with round-to-nearest-even then
matches a true fused f16 FMA exactly (24 significand bits >= 11 + 2).
Inf/NaN inputs turn the residual into NaN, the ordered compare skips the
parity fix, and IEEE special behaviour passes through unchanged. The
op_sel/op_sel_hi/neg_lo/neg_hi source modifiers apply to src2 through
the existing operand path; clamp stays rejected like the other packed
ops.

Every op in the 2Sum chain is decorated NoContraction: without it the
AMD RDNA3 Windows driver folds the sequence, collapses the residual to
zero, and the midpoint case decays to the double-rounded result. This
was caught by running the emitted shader on a real device (see below).

Verification:
- A mirror of the emitted sequence was checked against an exact
  integer reference (every finite f16 is m * 2^-24, so a*b + c is an
  exact Int128 multiple of 2^-48, rounded once to f16 RNE) across 34M
  cases: directed midpoint pins, random sweeps over all operand
  classes, tiny-addend midpoint stress, subnormal products, and
  Inf/NaN propagation. 0 mismatches.
- ShaderDump gains a pk-f16 program covering all five packed opcodes,
  both fma modifier paths, and the pinned constants; all programs
  decode and emit.
- The executable exec program now computes the pinned fma and its
  negated-addend twin (0x7A6B7A6B / 0x7A6A7A6A, straddling an f16
  midpoint) and stores them at offsets 20/24; GpuConformance checks
  both on device. All values match on an AMD Radeon RX 7700 XT.
2026-07-19 03:24:42 +03:00
Nicola Pomarico 09bd4f028b [Kernel] Implement sceKernelSyncOnAddressWait/Wake (#422)
libKernel's address-wait primitives were unimplemented, so every wait
returned immediately and guest runtimes that build spinlocks/queues on
top busy-spun forever. Implement them over the existing cooperative
block scheduler, keyed on the address, with a per-address wake
generation so a wait stays parked until a matching wake bumps it, and a
bounded self-heal deadline so a genuinely missed wake re-polls instead
of hanging.
2026-07-19 03:12:29 +03:00
ParantezTech 2bda253927 [script] added aerolib_catalog.py and docs/aerolib-catalog.md, renamed scripts/RELEASE-USE.md to docs/release-use.md 2026-07-19 01:33:15 +03:00
Berk 71e5912c75 [dotnet] remove lock files (#419) 2026-07-19 01:26:01 +03:00
Berk a030cb5a5d Gpu runtime stalls (#410)
* [runtime] restore default GC mode

* [cpu] add string leaf stubs

* [ampr] allow concurrent reads

* [bink] keep guest decode path

* [kernel] streamline host memory access

* [shader] add scalar memory fallback

* [gpu] bound guest data pool

* [gpu] reduce queue stalls

* [video] stabilize guest resources

* revert lock file
2026-07-19 00:31:50 +03:00
Dafenx 336286e588 CPU: scan final TLS access pattern offset (#414)
Co-authored-by: Dafenx <196083014+Dafenxz0@users.noreply.github.com>
2026-07-19 00:07:32 +03:00
Berk cab001f265 [GUI] Fixes click the controller B/O close button to close the game (#415) 2026-07-18 23:56:56 +03:00
Berk bab965e394 [HLE] Add RandomExports HLE (#413) 2026-07-18 23:44:57 +03:00
Spooks daaeb6213e Fix Massive Bug Preventing UE5 Titles From Booting (#406)
* Fix cross platform memcpy bug
2026-07-18 12:50:59 -06:00
Gutemberg Ribeiro 94153955b0 [Gpu] Metal backend: complete IGuestGpuBackend implementation on AppKit + Metal (#283)
* [ShaderCompiler.Metal] MSL translator core: dispatcher, EXEC model, compute stage

The Metal codegen backend, rebuilt on the merged backend-neutral
abstractions (replacing the pre-abstraction spike): consumes
(Gen5ShaderState, Gen5ShaderEvaluation) and emits MSL text; the renderer
owns MTLLibrary compilation, mirroring the emitters-produce-bytes rule
the Vulkan sibling documents.

The execution model mirrors Gen5SpirvTranslator: one invocation per GCN
lane (wave32 — natively the Apple simdgroup width), a typeless uint
register file with as_type<float> bitcasts, EXEC/VCC as per-lane bools
whose guest-visible mask registers materialize via simd_ballot, and the
same PC-dispatcher loop over basic blocks with the
SHARPEMU_SHADER_MAX_STEPS iteration guard and the dominating-scalar-
definition dataflow for buffer binding resolution. Unlike SPIR-V, MSL
permits shared prelude functions, so unaligned/subdword buffer access
is a range-checked device-uchar* helper instead of per-site inlining;
buffer byte lengths and the compute dispatch limit travel in one
reserved SharpEmuUniforms constant buffer (Metal has no OpArrayLength).

This first slice covers the compute entry point end to end: scalar/
vector ALU core (moves, int/float arithmetic, FMA family, shifts,
bitfield ops, min/max/med3, conversions, transcendentals with the Tau
scale on sin/cos), the full VCmp/VCmpx compare matrix writing VCC/EXEC,
the saveexec family, scalar compares and SCC-updating SOP2 forms, lane
ops (readfirstlane, mbcnt), VOP3 abs/neg/clamp/omod modifiers, scalar
memory, and raw global/buffer loads, stores, and atomics with EXEC
guards. Unsupported opcodes fail loudly with pc + mnemonic. SDWA/DPP,
typed format loads, LDS, images, and the pixel/vertex stages follow in
the next phases.

Tests live in their own self-contained project (the per-backend model:
depends only on the codegen under test): hand-assembled synthetic
fixtures drive the real decoder end to end, structural assertions and
golden-MSL comparisons run on every platform since translation is pure
text generation, and goldens regenerate via SHARPEMU_UPDATE_GOLDENS=1.

* [ShaderCompiler.Metal] Real-device runtime tests: compile + execute on the GPU

Lifts the spike's objc_msgSend LibraryImport harness (MTLDevice /
MTLCompileOptions with fast-math off, as a real Metal backend must
compile) onto the new translator contract: the guest data buffer binds
at index 0 and the SharpEmuUniforms constant buffer (dispatch limit +
buffer byte lengths) at index 1.

Three runtime tiers on hosts with a Metal device (no-op elsewhere so
Windows/Linux CI stays green): every fixture's emitted MSL must be
accepted by the OS runtime Metal compiler; the exec-store program must
produce bit-exact GPU results including the EXEC-masked store that must
not land; and a scalar countdown loop must iterate through the PC
dispatcher (s_cmp_lg_u32 + s_cbranch_scc1 across five round trips).
The loop fixture also fixes its own hand-assembly: s_sub_i32 sets SCC
to signed overflow, not result-nonzero, so the loop condition uses an
explicit compare.

* [ShaderCompiler.Metal] Phase 2: scalar/vector ALU parity with the SPIR-V translator

Ports the remaining ALU semantics from Gen5SpirvTranslator.Alu so the two
codegens cannot disagree on instruction behavior:

- Carry/borrow family (v_add_co/_ci, v_sub_co/_rev, v_subb/_rev) with the
  carry mask written to the VOP3 scalar destination or VCC, ANDed with
  EXEC; v_mad_u64_u32 with the 64-bit pair result and carry-out.
- Full SDWA support: byte/word source selects with sign-extension,
  integer abs/neg modifiers, and destination-select merge (zero-fill,
  sign-extend, preserve) into the previous register value.
- DPP16/DPP8: quad permute, row shl/shr/ror, mirror/half-mirror,
  broadcast and xor controls via simd_shuffle, bound-control and
  row/bank write-enable masks, fetch-inactive handling; DPP-predicated
  compares merge into VCC.
- Lane ops: readfirstlane from the first EXEC-active lane via
  ballot+ctz, readlane/writelane, permlane16/permlanex16.
- 64-bit scalar ops over SGPR pairs in real ulong arithmetic (logic
  family, shifts, bfe/bfm with width clamping, wqm quad expansion,
  cselect, mov, s_getpc) plus the B32/B64 saveexec families.
- Sopk forms decode the signed 16-bit immediate and s_cmpk compares the
  destination register; SOPC scalar compares including s_bitcmp0/1.
- Conversions: f16<->f32 via as_type<half>, pkrtz with round-to-zero
  mantissa truncation, pknorm via pack_float_to_{s,u}norm2x16, pk_u8
  byte insert, off_f32_i4 table, rpi/flr rounding; cube id/sc/tc/ma
  decision trees; v_cmp_class_f32; VCCZ/EXECZ/SCC readable as data.

Also fixes a real phase-1 bug the reference surfaced: fmamk/fmaak
sources arrive in natural order from the decoder, so all MAD/FMA forms
are fma(src0, src1, src2) — the previous operand swap computed
v1*v2+K for v_fmamk (should be v1*K+v2). The regenerated fmac golden
shows the corrected expansion, and mirroring the SPIR-V translator,
v_mul_u32_u24 is a full 32-bit multiply (only the hi/mad forms mask).

All 13 Metal tests pass including the real-GPU execution tier.

* [ShaderCompiler.Metal] Phase 3: typed format loads, LDS, and D16 subdword memory

Typed MUBUF/MTBUF loads convert through the descriptor's GFX10 unified
format at execution time, mirroring the SPIR-V translator: the prelude
bakes a 128-entry format table from the shared Gfx10UnifiedFormat
decoder (compiled shaders may be reused with new SRDs, so decoding must
stay dynamic), per-component layouts for the legacy DATA_FORMAT values
drive range-checked unaligned loads, NUM_FORMAT conversion handles
unorm/snorm (clamped at -1)/uscaled/sscaled/uint/sint/float including
f16 and the 10/11-bit unsigned mini-floats of 10_11_11 / 11_11_10, the
missing-component default is one in the format's domain, and dst_sel
swizzling comes from descriptor word 3. Format stores stay raw dword
stores like the reference.

LDS lands as 32 KB of threadgroup memory (gated on the program actually
using DS ops so occupancy is not taxed): ds_read/write b32/b64/b96/b128,
the write2/read2 pairs including st64 scaling, and ds_add_u32 as a
relaxed threadgroup atomic, with EXEC-guarded writes and the address
masked into bounds. Subdword loads/stores gain the D16/D16Hi variants
that merge into one half of the destination register (and shift the
source for high stores), classified the same way as the reference.

New GPU-executed fixture: an LDS round trip (write literal, s_barrier,
read back, store to the buffer) passes bit-exact on a real Metal device
alongside the existing tiers.

* [ShaderCompiler.Metal] Phase 4: pixel stage, images, and interpolation

The pixel entry points land with the same contract as the SPIR-V
translator (single-target and MRT forms, validated for unique guest
slots and dense host locations): the emitted fragment function takes a
stage_in struct carrying [[position]] plus the interpolated attributes
discovered from the program's V_INTERP controls, writes an output
struct with one [[color(hostLocation)]] attachment per binding typed by
its Float/Sint/Uint kind, seeds pixel-input VGPRs in SPI_PS_INPUT_ADDR
compact order from the fragment coordinate, keeps EXEC masking through
translation, and discards lanes that exit with EXEC off. Exports write
MRT targets per component under EXEC (disabled components keep their
previous value) including compressed half-pair exports; vertex-target
exports no-op until the vertex stage.

Images arrive as texture2d<float|int|uint> arguments (storage bindings
as access::read_write) with samplers alongside, classified from the
descriptor's unified format via the shared Gfx10UnifiedFormat decoder
and resolved per instruction with the same dominating-scalar-definition
scheme as buffers. The sample matrix covers implicit LOD, SampleL/Lz,
SampleB, SampleD gradients, PCF compare (manual reference<=texel,
broadcast r,r,r,1), per-lane texel offsets folded into normalized
coordinates by the selected mip extent (Metal sample offsets must be
constants), gather4 including compare and offset forms, clamped
ImageLoad/Mip, bounds-checked EXEC-guarded ImageStore, GetResinfo, and
A16 packed addresses / D16 packed data in both directions.

Graphics stages model LDS as per-invocation scratch (the SPIR-V
Private-array trick) instead of threadgroup memory. Wave ops keep the
invocation's real simdgroup in every stage — Apple fragment simdgroups
make that the same model as compute, where the SPIR-V translator
instead emulates a single logical lane; both round-trip EXEC masks
consistently.

The pixel fixture (interpolated attr0.xy plus inline constants exported
to MRT0) is golden-pinned, structurally asserted, and accepted by the
OS Metal compiler on a real device.

* [ShaderCompiler.Metal] Phase 5: vertex stage and fixed presenter shaders

The vertex entry point completes the four-entry-point contract: the
emitted vertex function takes fetched attributes as a stage_in struct
([[attribute(location)]], bound by the backend via MTLVertexDescriptor
from the reflected vertex inputs), returns [[position]] plus one
[[user(locnN)]] param output per export target 32..63 — unioned with
requiredVertexOutputCount so Metal's exact vertex-out/fragment-in
interface match succeeds, with unexported locations zero-filled —
seeds v5/v8 from [[vertex_id]]/[[instance_id]], intercepts buffer loads
the evaluator captured as fixed-function vertex inputs, and applies the
same EXEC-selected component rules to position/param exports (disabled
components default to 0,0,0,1) including compressed half pairs.

MSL vertex functions have no simdgroup attributes, so the vertex stage
models a single logical wave lane exactly like the SPIR-V translator's
graphics path: lane 0, ballot degrades to 0/1, and lane-shuffle ops
would fail Metal compilation loudly (no real guest vertex shader uses
them).

MslFixedShaders mirrors SpirvFixedShaders for the presenter surface:
the fullscreen-triangle vertex stage (position from the vertex index,
screen-space UV broadcast to every requested attribute location), the
copy/solid/attribute diagnostic fragments, and the output-free
depth-only fragment. Metal forbids "main", so each carries a stable
entry name.

The vertex fixture (constant position + one param export) is
golden-pinned and structurally asserted; it and all five fixed shaders
are accepted by the OS Metal compiler on a real device.

* [ShaderCompiler.Metal] Author static MSL blocks as template files

The prelude helpers (buffer access, ballot, tables), the format-load
conversion functions, and all five fixed presenter shaders move out of
AppendLine walls into Templates/*.msl embedded resources — real Metal
source with syntax highlighting and reviewable diffs — rendered by a
small {{placeholder}} substituter that fails loudly on any
unsubstituted token. Substitution points are deliberately few: the
stage-dependent ballot expression (vertex has no simdgroup attributes),
the baked GFX10 format table and layout cases, and the fixed shaders'
parameters. Per-instruction body emission stays programmatic, where a
template cannot express it.

Behavior-identical by construction: the golden files are untouched and
the whole suite — including the real-device execution and compile
tiers over the templated output — passes against them unchanged. The
.msl files carry no license headers (they would leak into every emitted
shader), so REUSE.toml annotates the Templates directory instead.

* [ShaderCompiler.Metal] Cover the MSL goldens in REUSE.toml

The golden files are verbatim emitter output regenerated by the test
suite; license headers inside them would either break the byte-exact
comparison or force the emitter to write SPDX text into every shader.
Annotate the Goldens directory like the Templates one.

* [ShaderCompiler.Metal] Address review: dominating-binding parity and harness binding indices

The buffer-binding fallback now mirrors the SPIR-V translator: a candidate
binding is accepted only when the descriptor registers hold the exact same
scalar definitions at the target PC as at one of the binding's own access
points (HasSameScalarDefinitions), instead of merely being non-conflicting
at the target. Resolutions are cached per PC like the reference.

The runtime test harness no longer hardcodes buffer indices 0/1 and a
20-byte uniforms blob: TryExecuteSingleThread takes the data/uniforms bind
indices, and ExecuteOrThrow derives them plus the uniforms size from the
compiled shader's GlobalMemoryBindings per the translator contract.

* [Gpu] Add the Metal guest-GPU backend: shader compilation and formats

First phase of the Metal backend behind the IGuestGpuBackend seam. The
backend compiles all three shader stages through Gen5MslTranslator and
exposes the guest render-target format table (mirroring the Vulkan table
case for case; guest format 9 maps to BGR10A2, the Metal layout matching
Vulkan's A2R10G10B10 pack). Wave64 compute is rejected with a clear error
until the two-pass emulation exists.

SHARPEMU_GPU_BACKEND=metal opts in on macOS; Vulkan stays the default on
every platform until the Metal presenter reaches parity. Presenter-side
methods fail loudly instead of dropping guest frames silently.

* [Gpu] Add the Metal presenter core: AppKit window, CAMetalLayer, CPU-frame path

The presenter opens an NSWindow hosting a CAMetalLayer and drives a manually
pumped NSApplication event loop, structured like the Vulkan presenter's
poll-and-render loop and posted onto HostMainThread the same way (AppKit
traps off the process main thread). All OS access goes through objc_msgSend
LibraryImport bindings declared locally — no windowing or binding packages on
this path, which is what keeps it NativeAOT-clean. Struct-returning ObjC
calls are avoided entirely so one calling convention works under Rosetta.

Presents CPU-produced BGRA frames and the splash through a fullscreen
triangle with a dedicated present fragment stage that flips V: with Metal's
y-up NDC the shared fullscreen triangle puts UV (0,0) at the bottom of the
screen while textures keep v=0 at the top. Frames letterbox via the viewport,
and nextDrawable paces the loop at presentation rate.

Guest-image submission now returns false (callers use their CPU-readback
fallback, which the presenter can show); draw and compute submission still
fail loudly pending later phases.

* [Gpu] Complete the guest-GPU seam: lift the AGC bypass surface onto the backend

The abstraction left AGC and VideoOut calling VulkanVideoPresenter statics
directly for guest work ordering (EnterGuestQueue, SubmitOrderedGuestAction,
SubmitOrderedGuestFlipWait, WaitForGuestWork), guest-image lifecycle (initial
data seeding, writes, fills, extents, upload tracking), the texture-content
cache probe, guest memory attachment, storage-offset alignment, perf
counters, and presenter close. With a non-Vulkan backend selected those
calls silently hit a never-started Vulkan presenter.

All of it now crosses IGuestGpuBackend: the Vulkan backend delegates to the
existing presenter statics (no behavior change), and the Metal backend
answers exactly like a presenter that is not running (sequence 0, image
unknown), which keeps callers on the same inline/CPU fallbacks they take
today. TextureContentIdentity moves to the seam types, and the bounded
AGC-to-presenter transfer pool becomes the backend-neutral GuestDataPool
(one pool by necessity: the AGC layer rents, the presenter returns).

AgcExports snapshots the backend's offset alignment once — it was a const
before and is read in per-draw loops (shader-key hashing, offset rounding).

* [Gpu] Metal guest work queue and guest images: ordered flips, writes, fills, blits

Mirrors the Vulkan presenter's execution model. AGC submissions become work
items consumed by the render loop in logical-guest-queue order: FIFO within
each guest queue, ready queues scheduled round-robin, completion tracked as
a contiguous sequence plus an out-of-order set, and producer backpressure
(count and payload caps) that consumer-enqueued follow-ups bypass to avoid
self-deadlock. The drain is budgeted (12ms, 256 items) so a backlog cannot
starve the Cocoa event pump or the present.

Guest images are Metal textures keyed by guest address, created on first
use from the registered display-buffer format tag (byte-identical encoding
to the Vulkan backend) and seeded once from pending initial data or guest
memory, since PS5 render targets alias guest memory. Coherence mirrors the
Vulkan design: DMA-style writes swap in a freshly written texture (never
mutating one an in-flight present may sample), fills clear through a
hazard-tracked render pass, and same-extent blits copy on the GPU.

Ordered flips capture the named image into an immutable version at their
exact queue position, so later work cannot change the frame a flip
selected; flip waits complete by queue position alone. Presentation picks
the newest ready queued guest frame (retiring superseded captures),
re-resolving mutable address-keyed textures at encode time so a write swap
never leaves a stale handle.

* [Gpu] Metal translated draws: pipelines, render state, bindings, write-back

Executes the seam's translated-draw surface on Metal. Offscreen, depth-only,
and storage draws are ordered guest work rendering into guest-addressed
images (published targets register as flip sources exactly like the Vulkan
backend); onscreen draws and recognized fixed-function draws ride the
presentation and render at present time into a pooled target.

Pipelines are built from GuestRenderState and cached by shader identity plus
a state hash: guest CB blend factor/op codes, write masks (bit-reversed for
MTLColorWriteMask), depth ZFUNC (bit-identical to MTLCompareFunction),
vertex attribute formats decoded from the same guest (dataFormat,
numberFormat) table the Vulkan backend uses, and RDNA 2:10:10:10 mapped to
Metal's R-low-bits 1010102 layout. Guest viewports pass through unchanged —
Metal accepts the negative heights PS5 games program, which is also how the
Vulkan backend inherits its orientation. Rect lists draw as 4-vertex strips;
Metal has no triangle fans, so those degrade to lists with a one-time warn.

Bindings follow the Gen5MslTranslator contract: global buffers at their flat
slot on both stages, SharpEmuUniforms (dispatch limit + buffer byte lengths)
after them, textures/samplers at the image slots with samplers decoded from
the raw guest descriptor words, and vertex streams at slot 26+ so they never
collide. Writable global buffers write back to guest memory before the work
item completes, preserving the CPU-visible GPU-write ordering point that
WaitForGuestWork promises. Feedback reads of a live render target sample a
blit snapshot; pooled guest data returns to GuestDataPool after upload.

Known simplifications for follow-up: textures upload a single mip level, and
the texture-content cache stays unclaimed (IsTextureContentCached=false)
until write-tracker-driven eviction exists, trading upload bandwidth for
correctness.

* [Gpu] Metal compute dispatch: the last seam gap

Guest compute dispatches are ordered guest work like draws. The uniforms
contract carries the per-axis dispatch limit (explicit thread counts when
the guest supplied them, groups x threadgroup size otherwise) so the
kernel's bounds guard clamps the overshoot threads of the last threadgroup;
threadgroup dimensions come from the translated shader, which bakes them at
compile time. Compute pipeline states cache per shader handle.

Storage images are shared live through the guest-image registry: a
dispatch's writes are visible to later draws, blits, and flips of the same
address, the address registers as a flip source at submit, and writer
sequences keep presentation waiting on exactly the work that produced the
frame. Writable buffers write back to guest memory before the work item
completes — the CPU-visible ordering point the returned sequence promises
through WaitForGuestWork.

Metal has no dispatch-base; nonzero base groups execute without the offset
behind a one-time warning until the emitted kernel grows base support.
SHARPEMU_SKIP_ALL_COMPUTE=1 skips all dispatches for hang isolation, same
as the Vulkan backend. With this the Metal backend implements the entire
IGuestGpuBackend surface — nothing throws.

* [Gpu] Address review: real bytes-per-pixel in guest-image uploads

Guest-image uploads hard-coded 4 bytes per texel, which mis-strided
Rgba16*/Rg32Float/Rgba32Float images and, in the guest-memory seed and
storage-snapshot paths, could make replaceRegion read past the managed
buffer. Texel width now comes from the pixel format, and
ReplaceTextureContents clamps the row count to what the source buffer
actually holds, so no caller can overread regardless of pitch and format.
RGBA8 initial data seeds only 4-byte-texel images; wider formats seed from
guest memory, whose layout is the image's native one. Extent byte counts
use the real texel width too.

Also restores the reference's comment on the deliberate single-item
backpressure admit: with no payload outstanding, refusing an oversized item
would wait forever since nothing is left to drain.

* [ShaderCompiler] First real-game fixes: SSendmsg no-op, scalar-state buffer declaration

Bring-up against a real title (2D engine, NGG shaders) found every draw
rejected at translation: RDNA2 NGG shaders bracket their exports with
s_sendmsg (GS_ALLOC_REQ/DEALLOC) to reserve hardware export space, and
neither translator handled the opcode — it fell through to the scalar-ALU
guard and failed with 'missing scalar destination'. Both translators now
treat SSendmsg as a no-op alongside SNop/SWaitcnt: exports are translated
directly, so the hardware message is moot. This was a shared gap, not a
backend one; the Vulkan path would reject the same shaders.

With translation unblocked, the OS Metal compiler rejected the emitted MSL:
the body reads the per-dispatch scalar-state buffer (initial SGPRs plus
per-binding byte biases) as b{initialScalarBufferIndex}, but the kernel
signature only declared the stage's own global bindings, so the name never
existed. The signature now declares it (const device — it is only read) at
its flat slot.

The presenter also logs one line when it first presents real content,
making 'window up but nothing shown' diagnosable from the log alone.
Verified: the title goes from 100% draw misses and a black screen to
~58k translated draws per minute and 4K frames presenting.

* [Gpu] Metal presenter: NSTimer-driven render loop under [NSApp run]

Replaces the hand-pumped event loop with a real running main loop. The
presenter now creates the NSApplication, orders the CAMetalLayer-backed
window on screen, and calls [NSApp run] so Core Animation's run-loop observer
actually commits presented drawables to the window server — without a running
loop the layer never composites and the window stays black regardless of what
is rendered into the drawable.

The per-frame work moves into RenderFrame, driven by a repeating NSTimer on
the main run loop (a tiny NSObject subclass whose onFrame: is an
UnmanagedCallersOnly callback, registered via the ObjC runtime — no binding
package). CADisplayLink is the natural choice and was tried first, but its
callback never fires in this process; proven in isolation against a bare
AppKit harness where a timer fires and composites and the display link does
not — the emulator runs as x86-64 under Rosetta and the display-server-backed
link is not serviced there. nextDrawable still blocks to the display, so the
timer only needs to keep up, not pace precisely.

Also fixes window sizing (the fixed 1280x720 window was being sized from the
guest 4K display mode, which macOS clamps while the layer keeps 4K geometry —
nothing visible), makes the metal layer the view's backing layer (wantsLayer
before setLayer) with an explicit frame, and stops both the AppKit loop and
the CFRunLoop on window close.

* [Gpu] Metal draws: normalize inverted viewports, resolve flips to drawn content

Two correctness fixes surfaced bringing a real title up. Guests program
Vulkan-style negative-height viewports for y-up rendering; Metal's NDC is
already y-up and rasterizes nothing for a negative height, so the viewport is
converted to the equivalent non-inverted rect (origin shifted, height
negated) with the same on-screen mapping.

Ordered flips now prefer produced content. A flip names the display buffer's
start address, but games render into the pixel surface past the buffer's
metadata block; the resolver takes the exact-address image when GPU work
wrote it, else the nearest same-extent GPU-written image within the buffer's
plausible metadata window, else the exact-address image even if only
seeded — so a flip presents the drawn frame rather than an empty seed. A
GpuWritten flag on guest images (set by draws, writes, blits, and dispatch
storage) distinguishes produced content from a speculative guest-memory
seed.

* [ShaderCompiler] Metal translator: per-stage uniforms slot, VCC/EXEC as data, exit branches

Three correctness fixes found by running a real game against the Vulkan
backend's behavior:

- Gen5MslShader carries UniformsBufferIndex: each stage emits its
  SharpEmuUniforms argument at globalBufferBase + totalGlobalBufferCount,
  and stages sharing a draw can disagree, so the presenter must bind the
  buffer per stage (Metal API validation: "missing Buffer binding at
  index 7 for sharpemu_uniforms").
- VCC (s106:s107) and EXEC (s126:s127) live in the scalar register file
  as raw 32-bit values with the per-lane bools as synced views. Programs
  legally park plain data in VCC (s_buffer_load into s[106] and then
  v_rcp_f32 of it); the bool-only model returned ballot masks instead.
- A branch to (or past) the program's end is an exit, matching the
  SPIR-V translator: sprite alpha-kill shaders use this to skip their
  tail and were rejected ("branch target outside program"), silently
  dropping every draw that used them.

Also: pixel-stage ballots use the per-lane form (this thread's own bit),
since simd_ballot is undefined inside the divergent dispatcher loop, and
v_readfirstlane returns the lane's own value under that model.

* [Gpu] Metal presenter: per-stage uniforms bind, keyboard input, perf overlay, title parity

- Bind SharpEmuUniforms at each stage's declared slot (see the paired
  translator change); one shared index left the vertex stage's slot
  unbound, zeroing its bounds-checked loads and killing interpolants.
- Vertex attribute byte offsets move onto the vertex descriptor (buffers
  bind at zero) and join the pipeline cache key, which they were silently
  missing from once baked into the pipeline.
- Unresolvable draw textures log a throttled warning instead of silently
  binding nothing.
- Keyboard input: an NSView subclass records keyDown/keyUp and feeds the
  POSIX host-input seam with a Windows-VK to macOS-keycode map covering
  the keys pad emulation polls; SHARPEMU_METAL_AUTOKEY scripts key
  presses for headless runs.
- Perf overlay (F1) drawn like the Vulkan presenter: CPU-rasterized
  panel uploaded to a small texture and composited with the present
  pipeline, with RecordPresent/RecordDraw feeding real numbers.
- Window title gains the selected GPU suffix and refreshes when the
  guest registers its application name; the layer is marked opaque so
  guest alpha never reaches the compositor.

* [Audio] Quiet sceAudioOutOutput on ports disposed by host shutdown

Closing the window disposes audio ports while guest audio threads are
still draining their last buffers; every remaining output then failed
the port lookup and logged a WARN per buffer (~190/s) until process
exit. Report success for missing ports once shutdown has begun; a bad
handle during normal operation still returns INVALID_ARGUMENT.

* [ShaderCompiler] Metal graphics stages model a single logical wave lane

The pixel stage used the real thread_index_in_simdgroup with all-ones
ballots while the vertex stage modeled lane 0 with 1-bit ballots, and
VReadlaneB32 still emitted a real simd_shuffle — reading another
fragment's register. Metal leaves simdgroup ops undefined inside the
divergent while(active){switch(pc)} dispatcher (empirically they
corrupted EXEC reconstruction), so graphics stages cannot use them.

Unify vertex and pixel on the SPIR-V translator's no-subgroup fallback:
one logical wave lane (lane 0), ballots degrade to bit 0, and the
shuffle-select family (readlane, readfirstlane, DPP16/DPP8 selects,
permlane16) resolves to the lane's own value. Writelane keeps the
lane-compare against the constant lane, matching the reference
fallback. Compute is untouched: its threads map one-to-one onto real
simdgroup lanes and still shuffle for real.

Entry parameter lists now always emit trailing commas and are closed by
one helper, so stage-specific trailing parameters no longer dictate
ordering.

* [ShaderCompiler] Metal compute mirrors the SPIR-V translator's wave semantics

Compute threads map one-to-one onto real simdgroup lanes, so restore
real simd_ballot for the compute prelude (the per-lane form was a
graphics fix that swept compute along) — masks parked in VCC/EXEC now
hold each lane's actual bit, and mbcnt/cndmask/saveexec read real
masks. VReadfirstlaneB32 broadcasts from the first guest-active lane
(ballot of EXEC, then ctz), matching the SPIR-V translator's explicit
first-active-lane broadcast rather than SPIR-V BroadcastFirst's
first-host-active semantics.

The wave64 gate moves into the translator and only rejects programs
that contain wave-sensitive operations (the SPIR-V translator's
subgroup-usage predicates: shuffle family, readfirstlane, wave control,
mbcnt, or VCC/EXEC operands). A wave64 kernel without them executes
identically per-thread on 32-wide Apple simdgroups, so it now
translates instead of being dropped.

* [Gpu] Metal draw textures resolve like the Vulkan presenter

Sampling a live guest target previously required the exact current
image at the descriptor's address; anything else silently bound
nothing. Mirror the Vulkan presenter's resolution chain:

- A descriptor naming a guest depth target's write or read address
  samples the depth image (identity channel select). Depth32Float
  cannot blit to a color format, so the ordered snapshot round-trips
  through a private staging buffer into an R32Float texture.
- Replacing a render target at the same guest address (new extent or
  format) retires the old image into a bounded variant cache instead of
  releasing it, and resolution scores the current image plus variants
  by descriptor match — exact extent over view format over
  initialization, active image breaking ties. Larger images qualify
  only for tiled descriptors, matching IsCompatibleGuestImageAlias.
- The throttled unresolved-texture warning remains the detector for
  anything the chain still cannot resolve.

* [ShaderCompiler] Document the Metal translator's wave-size model

Audit outcome for wave64 fidelity, no behavior change: every B64 mask
op, saveexec, and VCCZ/EXECZ test already reads and writes the full
register pair, lane indices never exceed 31 by construction, and the
GPU-executing runtime tests cover 64-bit exec save/restore. Record the
model in the class header.

* [Gpu] Plumb CB_BLEND constant color through both backends

The CONSTANT_COLOR / CONSTANT_ALPHA blend factors were mapped by both
backends but nothing ever supplied the constant, so any draw using them
blended against transparent black. Decode CB_BLEND_RED..ALPHA (the
constants existed unused) into GuestRenderState.BlendConstant, set it
dynamically per draw on both sides: Vulkan declares the blend-constants
dynamic state and calls CmdSetBlendConstants beside the viewport,
Metal calls setBlendColorRed:green:blue:alpha: in EncodeRenderState.

* [Gpu] Metal per-draw uploads bump-allocate from shared arena pages

Every draw created one MTLBuffer and one managed copy per binding
(padded guest globals, uniforms, vertex and index bytes), which
dominated allocation churn at hundreds of MB/s of garbage and held the
guest flip rate well under the display rate. Uploads now bump-allocate
256-aligned slices from 8 MiB shared-storage arena pages and bind by
offset; pages recycle once the last command buffer that referenced
them reports completion, polled at each drain so the ObjC interop
stays block-free.

Write-backs carry the slice's data pointer directly (the page outlives
the command buffer the caller waits on), the alignment-bias contract is
preserved by placing data at the bias inside its slice, and the padded
copies, per-draw buffer releases, and per-draw byte arrays are gone.

In-game on the test title: allocation rate ~795 to ~564 MB/s (the
remainder is the AGC-side per-draw guest snapshots), GC per stats
window ~30/30/17 to 16/16/15, CPU ~150 to ~131%. Metal validation
stays clean.

* [Gpu] Metal draw-texture cache: skip per-draw guest texel copies

Mirror the Vulkan presenter's identity-keyed texture cache: once the
render thread decodes a draw texture, the AGC submit thread skips the
guest-memory read/detile/copy for that identity entirely (the generic
IsTextureContentCached hook, which the Metal backend previously
hardcoded to false) and the render thread serves the cached MTLTexture
without re-uploading. GuestImageWriteTracker write-protects the source
pages; a guest CPU write evicts the entry at the next drain, and the
skip/eviction race self-heals by reading the texels directly.

Eviction differs from Vulkan in one deliberate way: dirty entries are
collected by address rather than identity, since ConsumeDirty clears
the flag on first read and several identities (same texels, different
samplers) can share one address.

Dreaming Sarah in-game on an M5 Max: guest flips 47 -> 60 (display
rate, matching Vulkan), ALLOC 564 -> 41 MB/s, gen0 GC 16 -> 5 per
second, CPU 131% -> 72%. Metal API validation clean; all 25 shader
compiler tests pass.

* [Gpu] Metal snapshot pool: recycle feedback-read textures and staging

Feedback reads created and destroyed an MTLTexture per draw (and for
depth sampling a private staging MTLBuffer too). Pool both with the
upload-arena lifecycle: acquisitions are tagged with the command buffer
that samples them at commit, and return to a bounded free list once it
reports completion. The command queue is serial, so the earlier
snapshot-blit command buffer is necessarily complete by then as well.

Dreaming Sarah renders correctly in-game; Metal API validation clean;
all 25 shader compiler tests pass. (The depth-sample path is exercised
only by inspection — no testable title samples depth yet.)

* [Gpu] Metal batched guest commands: one command buffer per drain

Draws and compute dispatches encode into a shared batch command buffer
committed once per drain instead of one commit per work item, mirroring
the Vulkan presenter's batched guest commands. Ordering inside the
batch is by encoder sequence: draw textures are now pre-resolved before
the consuming render or compute encoder opens, so feedback-read
snapshot blits encode into the batch (after the passes that rendered
the source) rather than committing ahead of them in separate command
buffers. Flips, image writes/blits, ordered actions, CPU-visible
write-backs, and every drain exit flush the batch first, preserving
the serial-queue ordering and WaitForGuestWork contracts.

Dreaming Sarah in-game on an M5 Max: CPU 72% -> 59% at a steady 60
guest flips; Metal API validation clean; all 25 shader compiler tests
pass.

* [Gpu] Metal vertex streams: share buffer slots, reject overflow gracefully

void Terrarium aborts with '-[MTLVertexAttributeDescriptorInternal
setBufferIndex:]: buffer index (31) must be < 31': every vertex
attribute got its own buffer slot from base 26, so six streams walk
past Metal's last vertex-stage buffer index (30) and the framework
assertion kills the process (reported by vladdenisov on PR #283).

Attributes of an interleaved vertex arrive from AGC as one stream
each, all reading the same guest buffer — assign slots by unique
(base address, stride, length) so those share one slot and one
upload. A draw whose unique streams still overflow the range is
skipped with a throttled warning instead of aborting. The assigned
slot keys the pipeline cache alongside the attribute offset, since
aliasing changes the baked vertex descriptor.

Dreaming Sarah renders correctly in-game at 60 flips with Metal API
validation clean; all 25 shader compiler tests pass. (void Terrarium
itself is not testable here — no decrypted copy.)

* [Gpu] Metal: drain guest work on enqueue, not only at render ticks

The Vulkan presenter's render loop is pulsed when guest work arrives
and waits at most a few milliseconds; the Metal render loop drained
guest work only inside its NSTimer tick, so every guest submit-then-
wait round-trip (release-mem labels, event writes, CPU-visible write-
backs) cost up to a full frame interval. Games that chain several such
waits per frame crawl: void Terrarium ran at 14 guest flips against
Vulkan's display rate, and input-to-effect latency suffered everywhere.

Enqueueing guest work now schedules a coalesced onGuestWork: message
onto the main run loop via performSelectorOnMainThread (block-free,
matching the NSTimer trampoline pattern), which drains the queue
immediately. A producer blocked on a full queue schedules the same
wake before waiting. void Terrarium's title menu: 14 -> 59 flips/s;
Dreaming Sarah unchanged at 60 with validation clean.

* [ShaderCompiler] Metal samplers: per-stage compact slots, not texture slots

Sampler argument indices copied the global texture slot (image binding
base + index), but Metal exposes only 16 sampler slots per stage
against 31 texture slots — a draw whose stages sample more than 16
images total emitted [[sampler(16+)]] and the MSL failed to compile
('sampler attribute parameter is out of bounds'), dropping the draw
(void Terrarium's in-game scenes).

Samplers now count sampled (non-storage) images from zero within each
stage, and Gen5MslShader carries the image-index -> sampler-slot map
plus the stage's image binding base so the presenter binds each
stage's samplers exactly where its shader declared them. A stage that
samples more than 16 images fails translation loudly. All 25 shader
compiler tests pass; goldens unchanged (single-texture fixtures keep
sampler 0).

* [Gpu] Metal draw textures: native guest formats, BC blocks, channel select

The draw-texture path assumed every texture was RGBA8: created
Rgba8Unorm, uploaded 4 bytes per pixel, and rejected anything whose
texel copy was smaller than W*H*4 as undersized. Games shipping
BC-compressed atlases (void Terrarium's entire in-game art) rendered
black, and because the rejected textures were never created they were
never content-cached — the AGC layer re-read and re-detiled megabytes
per draw (1.6 GB/s allocation, gen2 collections every second, 8 guest
flips).

Map guest texture formats to Metal case for case with the Vulkan
table (BC1-BC7 upload raw blocks — Mac-family GPUs sample them
natively — plus the 8/16/32-bit linear formats), size expectations
with the same block-aware byte math AGC uses, and honor the
descriptor's DST_SEL channel select through the texture swizzle,
mirroring Vulkan's component mapping. Unmapped codes keep the RGBA8
fallback.

void Terrarium now reaches gameplay past New Game: 49-54 guest
flips (from 8), no undersized-texture warnings, validation clean.
Dreaming Sarah unchanged at 60. All 25 shader compiler tests pass.

* [Gpu] Metal feedback reads: one snapshot per content version

Every draw sampling a live guest image blitted a fresh full-texture
snapshot, so compositing games that sample their render target on
most draws (void Terrarium: ~100 of ~105 draws per frame) pushed
gigabytes per second of blit traffic through the driver.

Guest images now carry a content version, bumped by every draw that
targets them, image write, blit destination, storage dispatch, and
guest-memory seed. The feedback-read path reuses one cached snapshot
until the version moves, so the blit happens per content change
instead of per draw. The image holds the snapshot's retain; consuming
command buffers keep replaced snapshots alive until they complete,
and retire/replace/write paths release the cache with the image.

void Terrarium in-game: 49 -> 58 guest flips at higher draw
throughput (Vulkan reference runs the same scene at 17-20 fps).
Dreaming Sarah unchanged at 60; validation clean; 25/25 tests pass.

* [Core] Pre-visit tracked texture pages before managed guest writes

A managed write into a page the guest-image write tracker has
protected dies with a fatal AccessViolation: the runtime surfaces
SIGSEGV in managed code as an exception before the resumable signal
bridge can restore access, unlike native guest stores which recover
through TryHandleWriteFault. Dead Cells crashed exactly there — an
AGC release-mem label write (CpuContext.TryWriteUInt64 on the render
thread) landing on a page the texture cache tracks.

TryWrite now calls GuestImageWriteTracker.NotifyManagedWrite up
front, unprotecting and dirtying any tracked pages in the span before
the copy — the hook existed for precisely this but had no callers.
Since this puts the tracker on every managed guest-write path, the
range snapshot now carries its overall bounds (one immutable object,
so the intersection test is always consistent with the array), letting
the common no-texture-pages case reject in a few instructions.

Dead Cells no longer crashes; Dreaming Sarah and void Terrarium
unaffected; all 25 shader compiler tests pass.

* [VideoOut] Name the active GPU backend in the macOS window title

macOS can run either backend — Vulkan through MoltenVK or native Metal
via SHARPEMU_GPU_BACKEND — so the window title now ends with the one in
use, e.g. "... · Apple M5 Max (Metal)" or "(Vulkan)". The suffix is
appended in SetSelectedGpuName (the single point both presenters call
to fold in the GPU name) and gated to macOS, so Windows and Linux
titles are unchanged. The name comes from a new BackendName on the
guest-GPU seam.

* [Gpu] Metal window: resizable, native full-screen, live drawable sizing

Add NSWindowStyleMaskResizable so the window can be dragged to any size
and set NSWindowCollectionBehaviorFullScreenPrimary so the green button
enters native full-screen instead of zooming. CAMetalLayer does not
track its drawable size to bounds on its own (even as a view's backing
layer), so the render loop matches drawableSize to the layer's current
bounds x contentsScale before each nextDrawable — a no-op on the common
unchanged tick. The present pass already aspect-fit letterboxes into the
drawable, so any window aspect ratio scales the frame without distortion.

Reading -bounds needs the x86-64 stret ABI for its 32-byte CGRect
return, added as SendStretRect. Verified live: drag-resize and
full-screen both scale correctly with Metal API validation clean.

* [ShaderCompiler] Metal wave64 compute: emulate cross-lane ops via scratch bridge

Replace the wave64 loud rejection with emulation, mirroring the SPIR-V
translator. A 64-lane guest wave is two 32-wide Apple simdgroups
co-resident in one threadgroup (Metal packs thread_index_in_threadgroup
0-31 into simdgroup 0, 32-63 into simdgroup 1), so sharpemu_lane becomes
thread_index_in_threadgroup & 63 and cross-lane ops that span the full
wave rendezvous the two halves through threadgroup scratch:

- ballot into EXEC/VCC/SGPR pairs: each half's simd_ballot is written to
  its scratch slot, a threadgroup_barrier syncs, and all lanes recombine
  the 64-bit mask into the low/high register pair (centralized in
  EmitBallotStore, which the wave32 path shares).
- read-first-lane: broadcasts the lowest active lane's value across both
  halves through a scratch slot (EmitWave64ReadFirstLane).
- mbcnt lo/hi: 64-lane thread-mask math (no cross-lane op, just correct
  per-lane masks; lanes >= 32 would overflow a 32-bit shift, so split).

The barriers are safe because the guest's scalar PC keeps all 64 lanes
lockstep through the dispatcher. Scope matches the SPIR-V reference: the
scratch is indexed by half, so correct for a one-wave (64-thread)
workgroup, and readlane across halves stays a 32-wide shuffle. Wave-
agnostic wave64 kernels still translate per-thread unchanged.

Verified on the real GPU (MetalRuntimeTests): the emitted wave64 MSL
compiles, and a 64-lane dispatch runs through the bridge barriers
without deadlocking, returning the broadcast value. All 27 tests pass;
Dreaming Sarah (60/60) and void Terrarium (in-game, 58 flips) show the
shared wave32 ballot path is unaffected.

* [Gpu] Metal samplers: bind through an argument buffer, lifting the 16-slot cap

Metal exposes only 16 direct [[sampler(N)]] slots per stage, but real
shaders sample more (void Terrarium's scene shader: 17 images) and were
dropped at translation. Route samplers through a per-stage argument
buffer instead: the MSL declares a Gen5Samplers struct (one sampler per
sampled image, [[id(N)]]) taken as constant& at a buffer slot past the
stage's globals/uniforms/scalar-state, and the runtime writes each
sampler's Tier 2 gpuResourceID into an arena slice bound there. Textures
stay on direct [[texture(N)]] slots (31 is enough). One sampler per
image keeps them distinct, matching the SPIR-V/Vulkan path — no dedup,
so no wrong-sampler artifacts.

Verified argument buffers lift the limit on Apple Silicon (20-sampler
pipeline probe). Dreaming Sarah renders correctly at 60/60 with Metal
API validation clean; the void Terrarium scene shader that exceeded the
limit now compiles and runs (draws 74 -> 102/frame); all 27 shader
compiler tests pass, goldens unchanged (fixtures sample nothing).

* [Gpu] Metal: Shared storage for CPU-populated, GPU-sampled textures

The MTLTextureDescriptor default is Managed, which on unified memory needs
an explicit host->device sync we never issue after replaceRegion, so the
GPU can sample stale texels. These textures are CPU-uploaded and GPU-read,
so Shared (coherent, no sync on Apple Silicon) is the correct mode.

* [HLE] Add missing AGC/AudioOut/Pad exports blocking Unity+FMOD titles

Four exports were unresolved and hard-stalled GPU/audio/input init in
Unity titles (Lunar Lander Beyond froze there before opening VideoOut):

- sceAgcDriverSetTFRing / sceAgcDriverSetHsOffchipParam: tessellation-ring
  and hull-shader off-chip config. We translate shaders directly, so these
  only need to report success for init to proceed.
- sceAudioOutGetPortState: report a connected primary output at full volume.
- scePadDeviceClassGetExtendedInformation: report a standard pad (no special
  peripheral) so device-class probes resolve.

Generic HLE, backend-agnostic (helps the Vulkan path equally).

* [VideoOut] RegisterBuffers2: mask the 32-bit category, accept COMPRESSED

sceVideoOutRegisterBuffers2's category is a 32-bit SceVideoOutBufferCategory
passed on the stack, but we read the full 64-bit slot — whose upper word
carries stale GNM magic (0xC0DEC0DE...) the caller never cleared. The old
check then rejected every call as INVALID_VALUE, so buffer registration
failed and no frame ever presented. Mask to 32 bits and accept both
UNCOMPRESSED (0) and COMPRESSED (1); we present either identically.

Fixes Lunar Lander Beyond reaching its window (now presents 3840x2160).

* [HLE] Stub sceAudioPropagation (3D-audio) so Astro Bot boots past its assert

Astro Bot hard-crashed right after the splash: it calls
sceAudioPropagationSystemQueryMemory during audio init, and because the
whole libSceAudioPropagation module was unimplemented the call failed, so
the game asserted (AudioPropagationContext.cpp:43) and executed int 0x41 to
abort — an unrecoverable trap that kills the process.

We don't model acoustic propagation (geometry-driven reverb/occlusion is a
quality feature, not a correctness gate). The API is placement-style, so
QueryMemory reports a buffer size and the rest succeed as no-ops: the system
lives in the caller's own buffer. All 39 entry points stubbed; the game now
boots past the assert to the presenter. Backend-agnostic HLE.

* [Kernel] pthread_cond_wait: don't spuriously EPERM an untracked mutex

pthread_cond_wait/timedwait required our host-side mutex tracking to show
the calling thread as the owner, else it returned EPERM. But libkernel's
uncontended mutex fast-path locks the mutex word in guest memory directly,
without an HLE call, so we often never observe the lock and see owner==0.

Real pthread_cond_wait requires the caller to hold the mutex but does not
verify it for normal mutexes, so EPERM here is doubly wrong: it spins the
guest (Hades hammered this millions of times/sec) and, worse, skips the
unlock — leaving the mutex held and wedging every thread that later blocks
on pthread_mutex_lock. When the mutex reads as untracked (owner==0), adopt
ownership so the unlock/wait/re-lock cycle is balanced and actually releases
it. Genuine ownership violations (owned by another thread) still error.

Eliminates the EPERM storm and converts the resulting livelock into correct
blocking; no effect on games that lock through the HLE (owner already set).

* [Core] SSE4a EXTRQ patch: read the xmm register from ModRM, not xmm2

The loader rewrites Sony's AMD-only SSE4a EXTRQ+blend idiom into SSE4.1 at
boot, because Rosetta 2 and Intel hosts raise #UD -> SIGILL on EXTRQ. The
matcher hard-coded the source register to xmm2 (ModRM 0xC2), but the compiler
allocates it freely: Dead Cells (PPSA15552) emits the identical idiom against
xmm1, so it slipped through unpatched and the game died with SIGILL right
after the first frame.

Read the register from the ModRM r/m field instead, covering xmm0-xmm7, and
require it to be consistent across the EXTRQ and the blend. The pure
match/encode logic is extracted into Sse4aExtrqBlendPatch, isolated from the
native page-patching, and unit-tested for every register plus the round trip
and rejection cases; DirectExecutionBackend just applies it.

Dead Cells now patches its xmm1 idioms and boots past the first frame.

* [Ngs2] Implement non-allocator sceNgs2SystemCreate / sceNgs2RackCreate

Dead Cells uses the non-allocator NGS2 create entry points, which were
unimplemented. sceNgs2SystemCreate came back as an unresolved import, so the
game got a garbage system handle; every downstream sceNgs2RackCreate /
sceNgs2RackGetVoiceHandle then failed, the voice handle stayed null, and once
gameplay started the audio path polled sceNgs2VoiceGetState/VoiceControl on
the null voice forever — freezing the game in-level at FLIP 0.

The non-allocator forms differ only in a caller buffer (rsi/rcx) vs an
allocator callback; the system/option and out-handle arguments sit at the
same positions, so they alias the existing WithAllocator implementations.
Resolves the NGS2 InvalidVoiceHandle storm (591+/run -> 0).

* [SaveData] Real save subsystem: ~/SharpEmu/Saves/<titleId>, events, full CRUD

Rework the SaveData HLE from a partial stub into a working subsystem:

- Storage moves to ~/SharpEmu/Saves/<titleId>/<dirName>/ (was next to the
  exe under user/savedata/<userId>/<titleId>), overridable via
  SHARPEMU_SAVEDATA_DIR. Metadata (title/subtitle/detail/userParam) and icon
  live in <slot>/sce_sys/. Pure path + param.json logic is isolated in a new
  SaveDataStorage type and unit-tested.
- Async event model: sceSaveDataGetEventResult now resolves (was an
  unresolved import a save worker polled forever), returning queued completion
  events or a clean 'no event' status; SyncSaveDataMemory posts a
  SAVE_DATA_MEMORY_SYNC_END event. Plus GetEventInfo/SetEventInfo/register
  callbacks.
- New exports: Mount/Mount2/Mount5/Umount, Delete/Delete5, GetParam/SetParam,
  SaveIcon/SaveIconByPath/LoadIcon, GetAllSize/GetProgress/GetMountInfo/
  IsMounted/GetSaveDataCount/GetMountedSaveDataCount/Abort, Initialize/
  Initialize2/Terminate, SaveDataMemory v1 aliases.
- Mounts are tracked so Umount2 really unregisters the /savedata0 mapping
  (new KernelMemoryCompatExports.UnregisterGuestPathMount) and params/icons
  resolve against the live mount; DirNameSearch surfaces param.json titles.

15 new unit tests (storage layout/sanitize/metadata + mount/event/param/delete
exports); full suite 277 passing.

* [Gpu] Metal: Cmd+F1 toggles Apple's Metal Performance HUD

Plain F1 keeps the built-in CPU-rasterized perf overlay; Cmd+F1 now toggles
the system Metal Performance HUD on the CAMetalLayer, Metal backend only.

Command-modified keys never reach keyDown: (AppKit routes them through the
key-equivalent chain), so the input view gains a performKeyEquivalent:
override that claims Cmd+F1 (also silencing the system beep) and leaves
everything else to the responder chain.

Configured per Apple's 'Customizing Metal Performance HUD':
developerHUDProperties with mode=default + logging=default, plus
MTL_HUD_LOG_SHADER_ENABLED=1 passed directly in the dictionary — HUD,
per-frame statistics logging, and shader-compile logging all enabled from
one property set; mode=disabled hides it again. Guarded by a
respondsToSelector: check for older macOS.

* [Gpu] Metal: also catch Cmd+F1 in keyDown: for the HUD toggle

Function keys reach keyDown: even with Command held (AppKit only reroutes
some chords through performKeyEquivalent:), so the HUD toggle was never
firing there. Handle Cmd+F1 in both the keyDown: and performKeyEquivalent:
paths, and keep it out of MetalHostInput so it can't also flip the plain-F1
perf overlay.

* [Audio] Diagnostics: NGS2 voice-param dump + AudioOut peak-amplitude trace

Two gated traces (idiomatic SHARPEMU_LOG_* style) that pinpoint where audio
dies for NGS2-based games:

- SHARPEMU_LOG_NGS2 now walks the sceNgs2VoiceControl param list and logs each
  {size,id} block header + payload bytes, confirming the real layout
  (header = u32 size, u32 id; waveform-block param id=0x10000001 carries the
  guest PCM pointer at +8; rate param id=0x10000005 carries the resample ratio).
- SHARPEMU_LOG_AUDIO_OUT logs sceAudioOutOutput call count and the peak
  amplitude of each submitted buffer.

Finding on void Terrarium: sceAudioOutOutput is called thousands of times on
both 8ch/float32 ports, but every buffer has peak=0.0 — the guest submits pure
silence. The host path (AudioOut -> PCM convert -> CoreAudio) is proven correct;
the silence originates in Ngs2SystemRender, which zeroes the output buffer
instead of mixing voices. Restoring audio for NGS2 games requires a real NGS2
software mixer (next).

* [Audio] NGS2 software mixer: decode + mix PS-ADPCM voices

NGS2-based games were silent because sceNgs2SystemRender only zeroed the
output buffer. This adds a real software mixer:

- Ngs2VagDecoder: clean-room PS-ADPCM ("VAGp") decoder producing mono PCM16
  with loop points resolved from the exact per-frame flag values (3=loop
  start, 6=loop end, 1/7=one-shot end).
- Voice control now parses the SceNgs2VoiceParamHead command list, decodes the
  waveform-blocks param's VAGp container once, and arms the voice.
- sceNgs2SystemRender mixes every armed voice belonging to the system into the
  leading grain of the render buffer as interleaved float32 (nearest-sample
  resample from the source rate to 48 kHz, additive into the front L/R pair),
  which is exactly what games copy to sceAudioOutOutput.

Verified on void Terrarium: previously peak=0.0 silence at AudioOut, now real
audible SFX/music. Voices are still armed on waveform assignment rather than an
explicit kick, so pooled/duplicate voices can overlap — trigger-state handling
is a follow-up.

* [Gpu] AGC: latch GPU-wait satisfaction to the produced value

Fixes a lost-wakeup race that stalled games at a black/splash screen. When a
RELEASE_MEM packet writes a completion label, the guest frequently resets that
label to 0 immediately to reuse it next frame. Our wake path
(GpuWaitRegistry.CollectSatisfied) re-reads *current* guest memory, so if the
reset lands before the wake pass runs, the transient satisfied window is missed
and the suspended DCB waits forever — even though the producing write executed
(traced as wrote=True) and its producer is marked completed.

RELEASE_MEM producers now call GpuWaitRegistry.LatchSatisfiedByValue with the
value they actually wrote, recording satisfaction at the moment of the write for
any waiter that value satisfies. CollectSatisfied honors the latch regardless of
the current (possibly-reset) memory value. This is fail-closed: a waiter only
latches when a real producer wrote a genuinely satisfying value.

Verified: Astro Bot's DEADBEEF sentinel wait (dcb.graphics waiting on a
release_mem label) that was permanently stuck is now resolved; void Terrarium is
unregressed (runs, audio intact, no producerless stalls). Astro still has
separate unresolved blockers (producer-behind-its-own-wait cascades and
producer=none-observed labels) tracked for follow-up; WRITE_DATA/DMA_DATA
producers could latch too but are left out until there is evidence they race.

* [Gpu] AGC: retry indirect dispatches whose GPU-computed dims aren't ready

GPU-driven games (Astro Bot) build their frame on the GPU: a compute dispatch
writes the thread-group dimensions for the next DISPATCH_INDIRECT into a guest
buffer. Our AGC parser reads those dimensions on the CPU at parse time, which
runs before the producing dispatch has executed on the render thread — so it
read 0/0/0 and dropped the work (agc.dispatch_reject zero-dimension), leaving
the scene unrendered (black) and cascading into stuck cross-queue fence waits.

Instead of dropping a zero-dimension INDIRECT dispatch, suspend the DCB on its
dimensions buffer (reusing the WAIT_REG_MEM suspend/resume + GpuWaitRegistry
machinery) until the producer writes non-zero dims, then re-parse and dispatch.
A bounded per-wait deadline (150 ms) resumes-and-drops a genuinely empty
indirect dispatch so it can never stall the queue, making the change
non-regressive: worst case matches the old drop behavior after a short wait.
Direct dispatches (dims inline) are unaffected.

Result: Astro Bot goes from a permanent black screen to actually rendering
(the presenter reports "Metal VideoOut presenting 3840x2160"). void Terrarium —
which issues no indirect dispatches — is unregressed (runs, audio intact, zero
rejects). Astro then hits a separate, newly-reached downstream crash (guest
TBB worker thread_set_state failure) tracked for follow-up.

* [ShaderCompiler] Metal: keep compute shaders within read_write and LDS limits

Two Metal limits made real Astro Bot compute shaders fail to compile/create,
which dropped their dispatches and cascaded into stuck GPU waits (splash hang):

- Textures with access::read_write are capped at 8 per function, but every
  storage image was declared read_write. Track each binding's actual access
  during body emission (ImageLoad->read, ImageStore->write, ImageAtomic and a
  load+store sharing one binding->read_write) and emit the minimal qualifier,
  so read-only/write-only storage images no longer count against the cap.

- Threadgroup memory is capped at 32 KB. A shader requesting the full 32 KB of
  LDS plus the separate 3-dword wave64 bridge overflowed by 12 bytes. Alias the
  bridge into the top of the LDS allocation when both are used, mirroring the
  SPIR-V translator's _waveScratchInLds path, keeping the total at 32 KB.

Verified on Astro Bot: "read_write access exceeds maximum (8)" and "Threadgroup
memory size (32780) exceeds maximum (32768)" are both gone; the 27 MSL golden
tests still pass (no golden used a storage image or LDS+wave64 shader).

* [Gpu] AGC: break cross-queue GPU wait deadlocks with a produced-value fallback

Real GPU-driven titles (Astro Bot) drive graphics and compute queues with
mutually dependent WAIT_REG_MEM fences: graphics waits on a compute EOP label,
compute waits on a graphics label. On hardware the two queues run concurrently
so the cycle resolves, but our submission parser is serial, so a label that gets
written -> reset for reuse -> re-waited across queues can wedge forever. The
latch fix helped the write-then-consume race but the cycle re-formed each frame
(graphics stuck at 3 flips, compute queues permanently suspended).

Producers now record the last value they wrote to each label
(GpuWaitRegistry.RecordProduced). A new deadlock breaker
(CollectDeadlockBroken, run from DrainResumableDcbs) releases any waiter stuck
past a 500 ms deadline whose condition is satisfied by that recorded value —
i.e. a real producer signalled the label at least once, guest memory has just
since been reset. It never fabricates a value, and the long deadline means
legitimate fences (which complete within a frame) never trip it.

Verified: Astro Bot goes from 3 flips (wedged on splash) to 25, loads its
splash level ("LevelDocument Loaded: ps_logo") and produces 2432x1368 frame
content. void Terrarium is untouched — 0 deadlock-break events, 1020 flips,
audio intact (its waits resolve far under the deadline). Tunable via
SHARPEMU_GPU_DEADLOCK_BREAK_MS.

* [Cpu] SSE4a EXTRQ patch: cover any blend destination register, not just xmm0

The EXTRQ+VPBLENDD idiom rewrite only matched when the blend destination was
xmm0 (VEX.vvvv byte 0x79). Sony's toolchain allocates that register freely: a
Dead Cells build emits `EXTRQ xmm4,0x28,0x00 ; VPBLENDD xmm3,xmm3,xmm4,2`
(VEX byte 0x61, dest xmm3). That instance stayed unpatched, so the AMD-only
EXTRQ reached Rosetta 2 and raised #UD -> SIGILL (0xC000001D) the moment the
game entered gameplay (loading level PrisonStart).

Read the destination register from VPBLENDD's VEX.vvvv / ModRM.reg as well as
the source from the ModRM r/m field, and emit PINSRD into that destination. Both
are still constrained to xmm0-xmm7 by the fixed VEX prefix. Match/encode stay in
the unit-tested helper.

Verified: Dead Cells now patches 14 EXTRQ blends (previously 0 on this build),
no SIGILL, and reaches PrisonStart. 15 patch unit tests pass, including the exact
xmm3/xmm4 bytes that faulted.

* [HLE] Implement Dead Cells' remaining unresolved imports

Three imports Dead Cells calls during boot/level-load were unresolved, so they
returned no defined value:
- scePadGetHandle (libScePad): returns the primary pad's handle (polled every
  frame for input); same validation as scePadOpen.
- sceNpEntitlementAccessGetAddcontEntitlementInfo (libSceNpEntitlementAccess):
  singular add-on-content lookup; we own no DLC, so zero the info out and return
  OK, matching the existing list variant.
- sceNpUniversalDataSystemEventPropertyArraySetString: telemetry setter, dropped.

Dead Cells now boots with zero unresolved imports. (It still stalls later at
PrisonStart level-load — a separate GPU/threading issue, not an import gap.)

* [Kernel] Fix pthread mutex deadlock: trylock semantics + stale-waiter clog

Hades hard-froze during boot on a "free but reserved" mutex: owner==0 yet
every acquisition failed forever. Two independent defects in the pthread
mutex compat layer combined to wedge it, both traced from real runs.

1. trylock incorrectly required an empty wait queue. POSIX
   pthread_mutex_trylock succeeds whenever the mutex is not currently held
   and owes no fairness to queued waiters; gating it on Waiters.Count==0 made
   a spin-on-trylock loop (which the game runs) spin forever against a single
   undrainable waiter even though owner==0. trylock now acquires on owner==0;
   the blocking lock still honours FIFO so genuine blocked waiters are not
   starved by a barging locker.

2. cond_timedwait timeouts leaked mutex re-acquire waiters. A cond wait's
   timeout enqueues a re-acquire waiter whose wake hand-off can be lost,
   orphaning it in the mutex queue. Multiple orphans from one thread piled at
   the FIFO head; the unlock hand-off then woke a dead wake-key and the mutex
   never drained. A thread can hold at most one pending acquisition on a
   mutex, so EnqueueMutexWaiterLocked now prunes any prior waiter for the same
   thread before enqueueing — collapsing the leaked pile.

Verified: Hades advances from a hard freeze at ~4.9M HLE calls (main and a
worker both blocked on the same free mutex) to 24.9M calls with no stall,
reaching the save-data/user-service boot stage. void tRrLM behaves
identically with and without the change (no regression); all 268 Libs tests
pass.
2026-07-18 20:32:00 +03:00
kostyaff 6dda6589d0 test: add Kernel/Loader unit tests (22 tests) (#373)
- SelfLoader: reject unknown magic, truncated headers; parse PS5 SELF embedded ELF
- KernelMemory: MapNamedFlexibleMemory/mprotect/munmap argument validation
- KernelEventQueue: create/delete/add/trigger/wait lifecycle

Co-authored-by: OMP <omp@local>
2026-07-18 20:19:55 +03:00
Aurélien Vivet 2ced3af114 AppContent: stub sceAppContentDownloadDataGetAvailableSpaceKb (#398)
Download data is not emulated as a real quota, so report a fixed 1 GiB
of free space and let titles skip the "storage full" path.
2026-07-18 18:44:30 +03:00
Berk 18708aa2d3 [GUI] Fixes and improvements for the GUI, including new image assets and updates to language files. (#400) 2026-07-18 17:44:04 +03:00
Berk a709ccca17 [shader_recompiler] Fix guest image byte count calculation for Vulkan video presenter (#395) 2026-07-18 16:09:26 +03:00
kadu04t 5309f384cf Reject undefined numeric LogLevel values (#390) 2026-07-18 15:47:19 +03:00
Mehmed Sinan Kömek e6be48a390 GUI: harden cross-platform updater integrity and rollback (#389)
* GUI: verify updater releases by commit and SHA-256

* GUI: add updater rollback and version safeguards
2026-07-18 14:41:59 +03:00
Raiyan b3e3fe5ea8 docs: note Windows on ARM runs the x64 build via emulation (#386)
Mirror the existing Rosetta 2 note for Apple Silicon: Windows on ARM
devices (e.g. Snapdragon) can run the Windows x64 build through Windows'
built-in x64 emulation.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-18 14:06:19 +03:00
Dmitriy c9e2a1a390 Update Russian translations (#387) 2026-07-18 13:52:08 +03:00
Granthik 01b80fe381 UPDATE: Harden bink2 bridge (#385)
fix stride overflow, check decode return values, null movie on open failure
2026-07-18 13:49:59 +03:00
999sian f84d869795 [Kernel] Match path cache comparisons to host filesystem case sensitivity (#381)
The negative-stat cache and the apr file-size cache memoize host
filesystem probe outcomes, but both were keyed with an ignore-case
comparer while the probes themselves (File.Exists/Directory.Exists/
FileInfo) are case-sensitive on Linux. That aliases distinct paths:

- stat("/app0/DATA.BIN") fails, the miss is cached, and a later
  stat("/app0/Data.bin") is answered NOT_FOUND from the cache without
  ever probing the disk - even though the file exists and the probe
  would succeed.
- sceKernelAprResolveFilepathsToIdsAndFileSizes serves the cached size
  of a case-distinct sibling file instead of the file's own size.

The registered-mount containment guard had the inverse problem: the
ignore-case StartsWith accepted a ".." path that resolves into a
sibling directory differing from the mount root only by case
("…/Save" vs "…/save"), letting guest I/O escape the mount.

All three sites now compare with the host filesystem's semantics:
ordinal-ignore-case on Windows, ordinal elsewhere. Windows behavior is
unchanged. Tests probe actual host filesystem behavior with real temp
files and skip their case-specific sections on case-insensitive hosts.
2026-07-18 13:36:01 +03:00
Spooks 13269797bf Add live debugger frontend and mutex stall recovery (#383) 2026-07-17 22:41:07 -06:00
jimmyjumbo 1c8cdd6537 [VideoOut] Initialize output options storage (#315)
* [VideoOut] Initialize output options storage

* [VideoOut] Keep output options size local
2026-07-18 04:14:27 +03:00
Peter Bonanni b566444df3 Add elapsed time to performance overlay (#250) 2026-07-18 03:35:56 +03:00
Raiyan 8a6f4f7826 [GUI] Per-game launch settings + shared SettingRow (#2) (#378)
Add per-game launch overrides (log level, import-trace limit, strict dynlib
resolution, log-to-file, and SHARPEMU_* environment toggles) with three-tier
resolution (per-game override -> global preference -> built-in default), stored
one file per game at user/custom_configs/<titleId>.json. Editable from a new
"Game settings..." context-menu dialog.

Introduce a shared SettingRow control and adopt it across the Options page and
the per-game dialog so the two read as one app. Fully localized (reusing the
existing Options.* keys), with the actions pinned in the dialog footer.
2026-07-18 03:19:37 +03:00
Peter Bonanni 41c9b44a8a [AJM] Track registered codec instance lifecycle (#352) 2026-07-18 03:02:19 +03:00
Mees van den Kieboom 743fe5cc26 [ShaderCompiler/Vulkan] Match vertex input numeric types (#351)
Declare UINT and SINT vertex attributes with integer SPIR-V component types so shader interfaces match the Vulkan pipeline formats. Keep normalized, scaled, and floating-point formats on float inputs.

Signed-off-by: missatjuhvdk1 <177474143+missatjuhvdk1@users.noreply.github.com>
Co-authored-by: missatjuhvdk1 <177474143+missatjuhvdk1@users.noreply.github.com>
2026-07-18 03:01:01 +03:00
Peter Bonanni 0b1dea43e8 [CPU] Preserve blocked leaf import waiters (#350) 2026-07-18 02:59:13 +03:00
lilp c9d018db8e Vulkan: fix guest storage image and render-state handling (#332) 2026-07-18 02:54:52 +03:00
Chris b479dc0466 videoout: implement output support query (#269)
Co-authored-by: Chris Cheng <chris@appxtream.com>
2026-07-18 02:44:23 +03:00
Peter Bonanni 22bbb4e909 [AvPlayer] Resolve guest media within app0 (#347)
Handle project-relative file URIs through the guest app0 mount, including unambiguous case-insensitive lookup for case-sensitive hosts.

Reject host paths, traversal underflow, malformed or remote URIs, and symlink/reparse escapes; cover accepted app0 forms and sandbox boundaries with nonparallel tests.
2026-07-18 02:42:30 +03:00
Peter Bonanni 3c500d2cf0 [SystemService] Write notice skip flag as byte (#346)
The Gen5 caller supplies a one-byte flag. Preserve pointer and memory-fault behavior while writing only that byte, and cover a seeded guest-memory boundary that rejects the former four-byte write.
2026-07-18 02:41:38 +03:00
Peter Bonanni bcb0ebd991 [ShaderCompiler] Fix VReadlane scalar destination field (#344)
V_READLANE uses the gfx10 VOP3A vdst byte even though its result is scalar. Decode bits 0-7 and cover the public LLVM s5 and s101 encodings so the VOP3B sdst field cannot be confused with this opcode again.
2026-07-18 02:41:07 +03:00
Mees van den Kieboom ecbb0db9be [Kernel] Preserve socket descriptors after failed connect (#343)
Keep ownership of a socket descriptor with the guest when connect fails, and route generic close calls through the socket table. Add a deterministic regression test for the failure and close sequence.

Signed-off-by: missatjuhvdk1 <177474143+missatjuhvdk1@users.noreply.github.com>
Co-authored-by: missatjuhvdk1 <177474143+missatjuhvdk1@users.noreply.github.com>
2026-07-18 02:40:11 +03:00
Jose Olguin Lagos 81633f6d5a Validate synthetic SPIR-V in CI (#335) 2026-07-18 02:39:07 +03:00
Zaid Yousef cc290f860b [Kernel] Return largest available direct-memory span (#334) 2026-07-18 02:38:39 +03:00
Berk ff3ac0bba1 chore: bump version to 0.0.2-beta.3 (#345) 2026-07-17 19:01:35 +03:00
tensorcrush 847371d2de [AGC] Decode VOP3P and emit packed f16 arithmetic (first slice) (#145)
* [AGC] Decode VOP3P and emit packed f16 arithmetic (first slice)

On gfx10 the VOP3P family lives under its own 0b110011000 prefix (word0
top byte 0xCC), which the major-opcode switch currently routes to the
SMEM branch, so packed instructions were decoded as scalar memory ops
and emitted as silent no-ops. Intercept the exact 9-bit prefix ahead of
the switch, decode the five packed-f16 arithmetic opcodes with their
op_sel/op_sel_hi/neg_lo/neg_hi/clamp modifiers, and emit them as
UnpackHalf2x16 -> component-wise f32 vec2 ops -> PackHalf2x16 so no
Float16 capability is needed.

Bit layout and opcode numbers pinned to LLVM MC test encodings
(vop3p.s, gfx10_vop3p_literalv216.txt) and VOP3PInstructions.td.
Unsupported modifiers, packed constants and out-of-scope packed opcodes
fail with a clear error instead of emitting wrong results.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* [AGC] Make packed f16 exact and drop v_pk_fma_f16 (review response)

Address the FP16 correctness review on the VOP3P slice.

Replace GLSL UnpackHalf2x16/PackHalf2x16 with explicit integer f16<->f32
conversions (EmitHalfToFloat/EmitFloatToHalf): exact widening with subnormal
normalisation, and narrowing with round-to-nearest-even, overflow-to-Inf and
NaN/Inf handling. Their subnormal and rounding behaviour no longer depends on
implementation-defined float-controls modes.

With exact conversions, v_pk_add_f16 and v_pk_mul_f16 are bit-exact to a true
f16 op (f32 result rounds losslessly to f16; a f16 product fits in f32). Emit
v_pk_min_f16/v_pk_max_f16 as fminnum_like/fmaxnum_like (NaN operand returns the
other; ordered numeric compare) instead of GLSL FMin/FMax.

v_pk_fma_f16 now fails emission loudly: a fused f16 FMA rounds once, an f32
multiply-add then pack double-rounds (fma(0x4100,0x7522,0x04EA) is 0x7A6B fused
vs 0x7A6A via f32). Exact fused emulation is a planned follow-up slice.

ShaderDump gains an Expect model (Translates/DecodeFails/EmitFails) and packed
regressions: arith, non-default modifiers, and loud-failure pins for the fma
case above and for clamp.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: tensorcrush <tensorcrush@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Antigravity AI <antigravity@gemini.com>
2026-07-17 18:58:22 +03:00
Berk fbafd3f429 [GUI] Add native Vulkan host surface support and more (#337)
* [GUI] Add native Vulkan host surface support and more

* reuse

* [GUI] set width session menu
2026-07-17 18:57:10 +03:00
Mees van den Kieboom aa25f6e978 [Loader] Restore PS5 SELF header support (#342)
Accept both PS4 and PS5 SELF signatures without treating version, key type, or flags as layout markers. Add synthetic coverage for valid header variants and malformed structural fields.

Signed-off-by: missatjuhvdk1 <177474143+missatjuhvdk1@users.noreply.github.com>
Co-authored-by: missatjuhvdk1 <177474143+missatjuhvdk1@users.noreply.github.com>
2026-07-17 18:55:16 +03:00
Raiyan b9eee497aa [VideoOut] Scope the AMD integrated-GPU penalty to Windows (#339)
* [VideoOut] Prefer real integrated GPUs over software rasterizers (#325)

Penalize only AMD integrated GPUs (the #97 vkCreateGraphicsPipelines
crash) instead of all integrated devices, so Intel/Apple/Qualcomm iGPUs
outrank Cpu-type software rasterizers (Mesa lavapipe). Hoist
ScorePhysicalDevice to the outer class and add unit tests for the
ordering.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* [VideoOut] Scope AMD iGPU penalty to Windows via a penalty helper

Extract ComputeDevicePenalty (the value subtracted from a device's base
score) and gate the #97 AMD-integrated penalty on Windows only. Mesa RADV
on Linux (e.g. the Steam Deck's AMD APU) is a different, working driver
and should keep its full integrated score. Add a Steam Deck test case and
drop inline comments.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* [VideoOut] Move device-scoring helpers out of the const block

Relocate ScorePhysicalDevice and ComputeDevicePenalty below the leading
const cluster instead of splitting it, and trim the vendor-ID reference
comment to adapters an x86-64 host can realistically enumerate.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-17 18:48:09 +03:00
Jose Olguin Lagos a9a4366ef4 test: cover Gen5 decoder fail-closed boundaries (#336)
Reimplement the five synthetic regressions from codefusion-repo/sharpemu#3 and #8 on sharpemu/sharpemu main befd20f415.

Keep the contributor-owned harness and corpus discussed in sharpemu/sharpemu#36 out of scope; use only a focused local recording memory fixture.

Signed-off-by: Jose Olguin Lagos <hellocodefusion@gmail.com>
2026-07-17 18:45:26 +03:00
ParantezTech 76e8f789c4 [github] edit contribution link in pull request template 2026-07-17 18:25:34 +03:00
ParantezTech 6fce0e3055 [github] update PR template and REUSE.toml 2026-07-17 18:20:44 +03:00
Berk befd20f415 Update contribution guidelines to enforce pull request template (#331)
Emphasize the importance of using the pull request template and outline consequences for non-compliance.
2026-07-17 13:52:41 +03:00
kuba 24b82a7f1c ci: rebuild the website when a release is published (#312)
* ci: rebuild the website when a release is published

* ci: add REUSE license header to notify-site workflow
2026-07-17 05:36:59 +02:00
Berk a9a4f511f0 Create pull request template for contributions
Added a pull request template to guide contributors in the SharpEmu Emulator Project.
2026-07-17 05:17:18 +03:00
Berk a44ce34a0a Add release process documentation to RELEASE-USE.md
Document the release process for SharpEmu, including preparation and publishing steps.
2026-07-17 05:00:22 +03:00
Berk 41ec56f813 chore: bump version to 0.0.2-beta.2 (#310) 2026-07-17 04:51:09 +03:00
Berk aa32d4bad0 added a release script to automate the release process (#309) 2026-07-17 04:48:46 +03:00
Hadi Abdulrahman faf3a397a8 [VideoOut] Fall back to 8-bit RGBA for unknown pixel formats instead of silently failing (#294)
Changes MapPixelFormatToGuestTextureFormat to default to format 56 (8-bit
RGBA) when the game uses a pixel format not yet in the known list, with a
stderr warning that reports the exact format value for project issue reports.

Previously, unknown formats returned 0, which caused RegisterKnownDisplayBuffer
to skip registration entirely. The GPU backend then couldn't find the buffer
during flip, producing vk.flip_capture_failed, and some games later hit a
Debug.Assert in ExecuteOrderedGuestFlipWait.

The fallback produces wrong colors for the affected games but lets them render
and display output, which is strictly better than a black screen or access
violation crash. The pixel format is printed to stderr so developers can
identify it and add proper support.

Co-authored-by: meowman <haadii2005@gamil.com>
2026-07-17 04:35:57 +03:00
samto6 488b285ecb [SaveData] Implement save data memory2 exports (#297)
* [SaveData] Implement save data memory2 exports

Astro Bot calls sceSaveDataSetupSaveDataMemory2 during boot and asserts
and null-writes when it fails, so the missing import surfaces as a
named crash. This implements setup plus the companion get, set, and
sync operations that make it useful. The store is one zero-filled file
per user and title at sce_sdmemory/memory.dat under the save root,
readiness is the backing file's existence, and get, set, and sync
return MEMORY_NOT_READY before setup. Struct offsets follow the
publicly documented homebrew savedata headers.

* [SaveData] Write setup result before mutating the memory backing file
2026-07-17 04:35:00 +03:00
samto6 1a4a2902d4 [HLE] Stub ContentExport, Font, and Pad calls Astro Bot needs to boot (#298)
Astro Bot asserts and null-writes when a subsystem init call fails, so
each missing import surfaces as a named crash. This stubs the blockers
observed during bring-up: content export init, eight font calls, and
pad tilt correction. sceFontGetHorizontalLayout writes the same
invented geometry as sceFontGetRenderCharGlyphMetrics and the rest
report success. Together with save data memory2 these take the title
to its splash image and font glyph rendering path.
2026-07-17 04:34:48 +03:00
kuba f544146d6d AGC: reduce inactive trace and draw allocations (#308) 2026-07-17 04:25:26 +03:00
kuba dbd74654c4 Vulkan: trim disabled draw diagnostics overhead (#307) 2026-07-17 04:18:06 +03:00
kuba 28485b60e2 CPU: avoid continuation emitter closure allocations (#306) 2026-07-17 04:17:58 +03:00
999sian 6b1abc1a38 perf(shader): cache pixel export masks (#288) 2026-07-17 04:17:39 +03:00
kostyaff 7494792249 [Tests] Add VirtualMemory and PhysicalVirtualMemory edge case tests (#266)
VirtualMemory: 4 new tests covering unmapped address returns false,
zero-length operations, page-boundary access, and cross-gap failures.

PhysicalVirtualMemory: 5 new tests covering lazy commit on demand,
reserve-only GetPointer, unmapped GetPointer returns null, free-list
first-fit reuse, and coalescing both neighbours on middle-range free.

Total: 9 new tests, 18 passed (8 VirtualMemory + 5 PhysicalVirtualMemory
+ 5 GuestMemoryAllocator). No production code changes.
2026-07-17 03:13:38 +03:00
jimmyjumbo 3585519007 [Audio] Correct float PCM endpoint conversion (#291) 2026-07-17 03:13:06 +03:00
Berk 8ee0fe8e44 Add files via upload 2026-07-17 03:03:05 +03:00
Tell-Shanks 0755ca15f7 core: report what occupies a fixed-address allocation on failure (#278)
* Implement DescribeAddressForDiagnostics method

Added a method to describe the state of a memory address for diagnostics.

* core: enrich allocation failure exception with host region diagnostic
2026-07-17 02:58:31 +03:00
jimmyjumbo 74c70cc2c1 [CI] Make versioned releases immutable (#289) 2026-07-17 02:54:35 +03:00
Berk 73e0ebf446 Update README order (#303) 2026-07-17 02:47:53 +03:00
Berk 459ae7e3f7 Docs/readme images (#302)
* [readme] update showcase images

* [readme] update readme images and change version to 0.0.2
2026-07-17 02:44:28 +03:00
kuba 8f405caebe Vulkan: reduce guest buffer upload allocations (#301) 2026-07-17 02:36:39 +03:00
kuba dabf723b3e [Vulkan] Honor guest depth clear state (#290) 2026-07-17 02:29:53 +03:00
AlexC 2129a12684 [GUI] Last commit SHA (id) displayed on about section (#279)
* Latest commit info axamal structure

* Link latest commit to text

* Added localization for About>Last commit info

* Made the commit hash a button that redirects you to the commit in github

* Reorder about section

* Commit and update icon in about section to have consistency in section
2026-07-17 02:25:01 +03:00
kuba 2db1fae282 [Tests/HLE] Cover APR resolve, stat, and streaming flow (#272) 2026-07-16 19:57:45 +03:00
kuba 33f96252da [CPU] Reject context transfers to unmapped guest addresses (#273) 2026-07-16 19:57:34 +03:00
kuba f7981a7ed7 [Tests/CPU] Verify import trampoline volatile-state ABI (#274) 2026-07-16 19:57:23 +03:00
kuba e10efa3ae1 [HLE] Make guest printf formatting locale-invariant (#271) 2026-07-16 18:39:48 +02:00
Peter Bonanni 16a2131b67 [PlayGo] Reject authoritative unknown chunk loci (#242) 2026-07-16 19:22:33 +03:00
Mees van den Kieboom 9883a9445d [AGC] Implement RDNA2 buffer/image/DS 32-bit atomic instructions (#222)
Adds decode and SPIR-V translation for the missing MUBUF, MIMG and DS
atomic instructions in the Gen5 shader translator, generalizing the
existing BufferAtomicAdd path. Covers swap, cmpswap, add, sub,
smin/smax, umin/umax, and/or/xor, inc and dec, plus the DS RTN
variants. Image atomics go through OpImageTexelPointer on the storage
image binding.

Notable: DS_CMPST operand order (DATA0 = comparator, DATA1 = new
value) is reversed relative to buffer/image cmpswap, which a dedicated
test locks in. ATOMIC_INC/DEC are approximated with
OpAtomicIIncrement/IDecrement, exact for the common 0xFFFFFFFF clamp.

Verified with 9 new synthetic decoder and end-to-end SPIR-V tests
(part of the #36 test corpus effort); full suite passes 35/35.

Signed-off-by: missatjuhvdk1 <177474143+missatjuhvdk1@users.noreply.github.com>
Co-authored-by: missatjuhvdk1 <177474143+missatjuhvdk1@users.noreply.github.com>
2026-07-16 17:28:39 +03:00
Peter Bonanni f2d9051358 Fix macOS fixed-address allocation collisions (#246) 2026-07-16 16:45:08 +03:00
Dafenx 53d414e096 Bound Vulkan host buffer pool memory (#201)
Co-authored-by: Dafenx <196083014+Dafenxz0@users.noreply.github.com>
2026-07-16 16:44:29 +03:00
MikeyLITE69 52d2874fa8 cpu: emulate BMI1/BMI2/ABM instructions in software when the host lacks them (#249)
## What

The native backend runs guest code directly on the host CPU. When the host doesn't
implement a BMI1/BMI2/ABM instruction that the PS5's Zen 2 cores do, it raises #UD
(STATUS_ILLEGAL_INSTRUCTION). Today the vectored handler just logs the faulting bytes
and gives up, so the title dies.

This adds a software fallback. On an illegal-instruction fault we decode the opcode
with Iced (the decoder already used elsewhere in the backend), evaluate it against the
trapped register/memory state, write the result and flags back into the CONTEXT record,
step RIP past the instruction, and resume.

Instructions covered (32- and 64-bit): ANDN, BLSI, BLSMSK, BLSR, BEXTR, BZHI, TZCNT,
LZCNT, RORX, SARX, SHLX, SHRX, PDEP, PEXT.

## Why

Users on CPUs without these extensions currently can't get past code that uses them.
This is a generic fix (no game-specific hacks) that improves compatibility on older
hosts. MULX is intentionally left out for now — its dest_hi/dest_lo operand ordering
is easy to get subtly wrong, so I'd rather add it separately with its own tests.

## How it's structured

- `BmiInstructionEmulator` holds the pure bit/flag semantics with no dependency on the
  unsafe CONTEXT plumbing, so it can be unit-tested directly.
- `DirectExecutionBackend.IllegalInstruction.cs` is the thin unsafe adapter (decode →
  read operands → emulate → write back → advance RIP). Anything it doesn't fully model
  returns false and falls through to the existing diagnostics unchanged, so it can never
  mis-handle an opcode it doesn't recognize.
- One hook in `DirectExecutionBackend.Exceptions.cs`, next to the other TryRecover* calls.
- Emits a single one-time "emulating in software" log line, not per-instruction spam.

## How I verified

- Added xUnit tests covering every instruction in both widths plus the CF/ZF/SF/OF
  edge cases (src == 0, shift-count masking, index beyond operand width, etc.).
- Cross-checked all the expected values against an independent reference implementation
  written from the Intel/AMD definitions; results match.
- `dotnet build` + `dotnet test` pass locally.

## Notes

New files follow .editorconfig (4-space, SPDX headers, REUSE-compliant).
2026-07-16 16:38:27 +03:00
Gutemberg Ribeiro c5a82c1065 [Perf] Behavior-preserving hot-path allocation/LINQ wins (#264)
* [Perf] Gate event-flag tracing so it allocates nothing when disabled

TraceEventFlag built its interpolated argument (and, on the wait path,
FormatFrameChain + a new StringBuilder(256) in FormatGuestWaitObject plus
~12 guest-memory reads) on every call, then checked the env var inside the
method — so every sceKernelSetEventFlag/Clear/Poll/Wait paid a string
allocation and an Environment.GetEnvironmentVariable P/Invoke even with
tracing off. Cache the flag once in a static readonly bool and gate every
call site, matching the semaphore/event-queue pattern. Behavior is
unchanged when SHARPEMU_LOG_EVENT_FLAG=1.

* [Perf] Hoist IsNoBlockLeaf classification to import-stub setup

The leaf-dispatch path called IsNoBlockLeafImport(nid) — a ~30-literal
string pattern match — on every leaf import. IsLeaf/NidHash were already
precomputed on ImportStubEntry at stub setup; add IsNoBlockLeaf alongside
them and read the field in the hot path. Behavior unchanged.

* [Perf] Cache trace env-var flags read on hot paths

Two SHARPEMU_LOG_* env vars were read via Environment.GetEnvironmentVariable
(a P/Invoke + transient string) on hot paths: SHARPEMU_LOG_FIBER on every
fiber context transfer, and SHARPEMU_LOG_DIRECT_MEMORY on every direct-memory
op (~8 sites). Cache both once — _logFiber alongside the other backend _log*
flags, and _traceDirectMemory behind the existing ShouldTraceDirectMemory
helper. Behavior unchanged.

* [Perf] Avoid per-iteration thread snapshot in the idle pump loop

PumpUntilGuestThreadsIdle allocated a full GuestThreadState[] snapshot
(via LINQ Values.ToArray()) on every spin just to tally three run-state
booleans. Tally them under the lock with an allocation-free helper, and
only materialize the snapshot inside the gated (default-off) diagnostic
dump. SnapshotGuestThreads now uses Values.CopyTo instead of LINQ.
Behavior unchanged.

* [Perf] De-LINQ GetPixelColorExportMask on the per-draw path

GetPixelColorExportMask ran a Select/OfType/Where/Aggregate chain over all
shader instructions, allocating iterators + closures, and is called per
render target (twice per draw via CreateRenderState and HasPixelColorExport).
Replace with a manual scan producing the identical mask — no allocation, and
it removes an authored-LINQ use the repo bans.

* [Perf] De-LINQ per-draw render-target selection in AgcExports

The bound-render-target selection used Where(...).OrderBy(...).ToArray()
(plus a second Where/ToArray fallback) on every translated draw, allocating
LINQ iterators + closures. Replace with an explicit filter into a pre-sized
list + List.Sort by slot; slots are distinct so this matches the stable
OrderBy. Same result, no per-draw LINQ allocations.

* [Perf] De-LINQ per-draw render-target validation in the Vulkan presenter

SubmitOffscreenTranslatedDraw validated its render targets with
targets.Any(...) twice plus targets.Select(a=>a.Address).Distinct().Count()
(a per-draw HashSet), and broadcast a single blend with
Enumerable.Repeat(...).ToArray(). Targets are <= 8, so replace with manual
scans (invalid-target check; combined dimension-mismatch + pairwise
aliasing check) and Array.Fill. Same results, no per-draw LINQ allocations.

* [Perf] Binary-search VirtualQuery region lookup (SortedList)

_mappedRegions was an unordered Dictionary, so TryFindVirtualQueryRegionLocked
scanned every region for containment/next — O(n) per sceKernelVirtualQuery and
O(n^2) when an allocator walks the address space with the findNext flag. Store
regions in a SortedList keyed by base address (every write already uses the
region's own address as the key) and find the containing/next region with a
binary search over the sorted keys. Also drops a now-redundant Values.OrderBy.

Non-overlapping regions assumed (mmap semantics), so only the floor region can
contain the query. Behavior-preserving; worth spot-checking VirtualQuery-heavy
titles.

* [Perf] Remove stray CLI packages.lock.json committed by mistake

git add -A in an earlier commit swept in a regenerated
src/SharpEmu.CLI/packages.lock.json. main tracks no lock files (central
package management, no RestorePackagesWithLockFile), and REUSE.toml no
longer covers packages.lock.json, so the committed file failed the REUSE
Compliance check. Remove it.
2026-07-16 16:37:28 +03:00
can a8be1daca4 Fix sceNetHtonl/Htons/Ntohl/Ntohs discarding their converted result (#211)
Each of the four byte-order helpers computed the swap into Rax and then
returned via ctx.SetReturn(0). SetReturn writes its argument into Rax, so it
immediately overwrote the converted value with 0 and the guest saw every
sceNetHtonl / sceNetHtons / sceNetNtohl / sceNetNtohs call return 0.

Leave the swapped value in Rax and return ORBIS_GEN2_OK as the dispatch status
instead, matching how the value-returning exports in this file (e.g.
sceNetPoolCreate) already work.

Add a NetExports test suite covering the swaps, the 16-bit width masking, an
htonl/ntohl round-trip, and a regression guard that a non-zero input never
converts to 0.
2026-07-16 16:24:34 +03:00
999sian 26dbad8ac8 fix(input): map POSIX stick endpoints exactly (#244) 2026-07-16 16:19:33 +03:00
Dafenx 2196a1f786 Gate Vulkan texture upload traces (#220)
Co-authored-by: Dafenx <196083014+Dafenxz0@users.noreply.github.com>
2026-07-16 16:19:11 +03:00
kostyaff 8db2c0622d [VideoOut] Implement sceVideoOutIsOutputSupported stub (fixes #257) (#270)
Quake (PPSA01880, Kex Engine) crashes at startup because
sceVideoOutIsOutputSupported (NID Nv8c-Kb+DUM) is unimplemented.
The game calls it to check video output capabilities before
initializing rendering.

Add HLE export stub: returns 1 (supported) for SceVideoOutBusTypeMain,
0 otherwise. The emulator renders via Vulkan and supports any pixel
format or aspect ratio on the main bus.

NID verified via Ps5Nid.Compute SHA1 algorithm.
Export name confirmed in scripts/ps5_names.txt.
2026-07-16 16:07:21 +03:00
mskomek 53da00fa89 GUI: add cross-platform release updater (#243) 2026-07-16 14:01:42 +03:00
Spooks b4b95014f1 Fix/deadcells crash (#262)
* Boot compatibility fixes for UE titles, GUI toggles, and DeS render/boot work

Checkpoint of the Monster Truck Championship and Demon's Souls boot work.
Each piece is independently useful and verified against the titles.

- playgo: scePlayGoGetLocus now returns BAD_CHUNK_ID for chunk ids outside
  the known set, matching real firmware. Titles enumerate chunk ids until
  that error; answering OK for every id made the scan wrap the ushort range
  and spin forever. Missing-sidecar and no-app0 fallbacks report a
  fully-installed single chunk 0 so scePlayGoOpen keeps succeeding.
- kernel: restore the SHARPEMU_WRITABLE_APP0 opt-in. Unpackaged UE dumps
  write their Saved tree under /app0 during PS5 component init and treat
  the denial as a fatal boot error.
- pad: accept handle 0 as the primary pad across all pad calls. Real
  firmware hands out small non-negative handles and some titles read state
  with handle 0.
- bthid: env-gated experiment hooks for the Thrustmaster wheel middleware
  investigation (fail-only-RegisterCallback modes and a synthetic
  enumeration callback with a zeroed event struct). All default off.
- gui: add SHARPEMU_LOG_IO and SHARPEMU_WRITABLE_APP0 toggles to the
  Environment tab.
- videoout: per-swapchain-image render-finished semaphores (the shared
  semaphore raced the swapchain); whole-mip-chain layout init for offscreen
  guest images (sampled binds read mips stuck in Undefined); GPU-resident
  texture availability now canonicalizes through the texture format table
  and accepts compatibility-class aliases, cutting per-frame CPU texture
  re-reads (143 GB -> 55 GB per 300 s in Demon's Souls, 0.2 -> 0.5 fps).
- hle: add sceSystemServiceGetNoticeScreenSkipFlag,
  sceSystemServiceGetMainAppTitleId (title id published from the runtime),
  and sceNpWebApi2CreateUserContext (refuses so the online layer backs off).
- rtc: SHARPEMU_RTC_PROBE_RANGE diagnostic dumps the code around a busy-wait
  caller of sceRtcGetCurrentTick once; costs nothing when unset.

* [ShaderCompiler] Fix VReadlaneB32 scalar destination field

The scalar destination lives in the low vdst byte (bits 0-7); it was read
from bits 8-14, the VOP3B carry-out field readlane does not have, sending
every readlane result to s0. Verified against raw gfx10 encodings and
LLVM's assembler tests (v_readlane_b32 s5, v1, s2 -> low byte 0x05).

* [VideoOut] Survive device loss and flip-order asserts without dying

Two ways a frame could take down the whole presenter:

- Device loss between any two Vulkan calls in a frame unwound the window
  thread, and the Dispose-time fence check then threw again, masking the
  original error. Catch the loss at the frame boundary, retire
  presentations and guest submissions whose fences can never signal, and
  keep the window loop pumping so the game (audio, logic) carries on.
- The ordered-flip capture invariant is violated ~100 times per run by
  Demon's Souls (PPSA01342); on debug builds the Debug.Assert fail-fasts
  the process with nothing in the log. Downgrade it to a once-per-version
  warning until the capture/wait ordering is understood.

* Fix Dead Cells shader cache regression

---------

Co-authored-by: StealUrKill <35749471+StealUrKill@users.noreply.github.com>
2026-07-16 13:57:16 +03:00
ParantezTech 7b91c964fc [VulkanVideoPresenter] revert back icon 2026-07-16 04:51:25 +03:00
Spooks 9bacb883f1 Fix Linux aligned mapping retention (#247)
* Fix Linux aligned mapping retention

* Cover Linux aligned mapping retention
2026-07-15 19:34:35 -06:00
Miguel Cruz 864cbb0fa0 [AGC/Vulkan] Extend PS5 runtime and rendering compatibility (#216)
* [Core] Add POSIX native execution and PS5 SELF support

Extend the native backend, guest TLS, fixed-address memory, and loader paths needed by PS5 titles on Windows, Linux, and macOS. Keep workstation GC so high-core-count hosts do not reserve over fixed guest image bases.

* [HLE] Expand PS5 service and media compatibility

Add the kernel, threading, save-data, networking, audio, video-codec, font, dialog, and service exports required by newer PS5 software. Preserve every SysAbi NID currently registered by main while adding the compatibility surface used by ASTRO BOT.

* [AGC/Vulkan] Extend Gen5 shader and presentation support

Expand PM4 handling, Gen5 shader translation, MRT and packed export support, guest image tracking, depth initialization, texture aliasing, and Vulkan presentation. Add the performance overlay and address-filtered diagnostics used to validate ASTRO BOT with original shaders.

* [Core] Align static TLS reservation across hosts

* [Pad] Align primary user ID with UserService

* [Gpu] Preserve runtime scalar buffers across renderer seam

* [AGC] Restore omitted command helper exports

* [Vulkan] Reuse primary views for promoted MRT targets

* [Vulkan] Preserve scratch storage bindings in compute dispatches
2026-07-16 02:02:34 +03:00
Berk ad5a7d3799 [CI] fix asset (#245) 2026-07-16 01:05:28 +03:00
en-he f69fdd4027 Fix PNG chunk CRC validation (#241) 2026-07-16 00:46:43 +03:00
Cloudy 9eef39549a hle: implement offline NP reachability state (#238) 2026-07-16 00:35:18 +03:00
Nicola Pomarico 90651f26a9 Fix macOS exact-address mmap falling back to MAP_FIXED (#214)
The PS5 image loader requires guest images to land at fixed virtual
addresses (e.g. the 32 GiB main image base). On macOS the exact-address
mmap path only ever passed that address as a hint, never MAP_FIXED, to
avoid clobbering untracked host mappings (dyld, JIT heap, Rosetta).

On Apple Silicon under Rosetta 2 the kernel does not reliably honor
that hint, so the allocation failed deterministically before a single
guest instruction ran (reported in #194). Retry with true MAP_FIXED
as a fallback only when the hint-only attempt doesn't land at the
requested address, so the safer hint path is still tried first.
2026-07-16 00:28:39 +03:00
kostyaff b5930465f2 [AGC] Correct VReadlaneB32 decode and emission (#237)
VReadlaneB32 (VOP3 0x360) had two bugs:

1. Decode: VOP3 decode always set destinations = Vector(word & 0xFF).
   For VReadlaneB32, bits 0-7 are unused — the scalar destination is
   in bits 8-14. Now decodes as Scalar((word >> 8) & 0x7F).

2. Emission: VReadlaneB32 was in the VMovB32 fall-through group,
   just returning GetRawSource(instruction, 0) — reading the current
   lane's value. By ISA, sdst = vsrc0[lane(src1)], which requires
   reading a different lane's value. Now uses SPIR-V
   GroupNonUniformBroadcast(scope=Subgroup, value=src0, lane=src1)
   when subgroup operations are available, with a fallback to the
   current-lane simplification when not.

3. Routing: TryEmitVectorAlu called TryGetVectorDestination first,
   which checks for VectorRegister kind. With the scalar destination
   fix, VReadlaneB32 now routes to a new TryEmitReadlane handler
   before the vector destination check.

4. Subgroup capability: UsesSubgroupShuffle now includes VReadlaneB32
   so GroupNonUniform capability is enabled when needed.

Verified: dotnet build 0 errors/0 warnings, 26/26 tests pass,
ShaderDump all programs behaved as expected.
2026-07-16 00:27:26 +03:00
brbrhuehue-matrix 83303e953e Update Brazilian Portuguese translation (#233) 2026-07-16 00:26:08 +03:00
kostyaff 92497689ab [AGC] Fix VOP3 decode for V_READLANE_B32 and V_WRITELANE_B32 (#232)
PR #200 moved shader files from SharpEmu.Libs/Agc/ to new projects
SharpEmu.ShaderCompiler and SharpEmu.ShaderCompiler.Vulkan. This
re-ports the VOP3 decode fix from PR #226 to the new file paths.

Decode table (Gen5ShaderTranslator.cs):
- 0x360: VMadU32U16 -> VReadlaneB32 (per RDNA2 ISA)
- 0x361: VMulLoU32 -> VWritelaneB32 (0x361 was a duplicate of 0x169)
- 0x373: added VMadU32U16 at its correct opcode

Emission (Gen5SpirvTranslator.Alu.cs):
- VWritelaneB32: per-lane conditional write via IEqual+Select,
  stores with guardWithExec:false (writelane bypasses exec mask)
- VReadlaneB32: kept as GetRawSource(instruction, 0) simplification
  (correct emission with src1 lane select is a follow-up)

Verified: dotnet build 0 errors/0 warnings, 26/26 tests pass,
ShaderDump all programs behaved as expected.
2026-07-16 00:25:59 +03:00
Gutemberg Ribeiro de7973af55 [CI] Fix PR artifact-links comment for fork PRs (#227)
The workflow resolved the PR from workflow_run.pull_requests or the head
commit, both of which come back empty for PRs from forks — so it logged
"No open PR for this build" and never commented. Resolve the PR by its
head "owner:branch" (from workflow_run.head_repository/head_branch),
which works for forks, keeping the commit-association path as a fallback.
2026-07-16 00:25:21 +03:00
s1099 3da7ae7d70 Update repo URL (#224) 2026-07-16 00:24:14 +03:00
Jack Del Aguila 1be009ce40 [HLE] Fix POSIX condition variable semantics (#113) (#223)
* [HLE] Trigger AGC graphics events by filter instead of exact ident (#173)

The PM4 EVENT_WRITE packet carries a 6-bit hardware EVENT_TYPE, but the
guest registers AGC events via sceAgcDriverAddEqEvent with a full guest
eventId. These two values are not the same numbering scheme, so the exact
ident lookup in TriggerRegisteredEvents never matched and the AGC
interrupt thread hung forever.

Add TriggerRegisteredEventsByFilter, which wakes every graphics event
registration on every queue. This is a compatibility workaround for
issue #173 while the real PS5 mapping remains unknown.

Includes unit tests covering the mismatched ident/eventType case.

* [HLE] Fix POSIX condition variable semantics (#113)

Remove PendingSignals from PthreadCondState. POSIX condition signals are edges,
not semaphore credits - a signal with no waiter must have no effect. The previous
implementation persisted signals, causing lock inversions and predicate bypasses.

Changes:
- Remove PendingSignals property and TryConsumePendingSignal method
- Remove pending signal consumption logic from PthreadCondWaitCore
- Remove PendingSignals increment from PthreadCondSignalCore
- Add regression tests verifying POSIX-correct behavior

Fixes #113
2026-07-16 00:24:05 +03:00
999sian ad92ab30fd fix(gui): respect case-sensitive Linux paths (#218) 2026-07-16 00:17:31 +03:00
can 5f97d2edc4 Fix sceKernelGetTscFrequency disagreeing with the counter on non-Windows hosts (#213)
sceKernelReadTsc only returns the CPU's RDTSC when the host RDTSC reader is
available (currently 64-bit Windows); on Linux and macOS it falls back to the
QPC-based Stopwatch. ResolveKernelTscFrequency, however, still consulted the
CPUID-reported hardware TSC frequency in that case, so sceKernelGetTscFrequency
reported a multi-GHz rate while ReadTsc was ticking at the Stopwatch frequency.
A guest computing elapsed = readTscDelta / frequency then gets the wrong time
on those platforms.

Gate the calibrated/CPUID frequencies on RDTSC actually being available (the
calibration path was already self-gated; the CPUID path was not) and otherwise
report the Stopwatch frequency, keeping ReadTsc and GetTscFrequency consistent.

The selection logic is extracted into a pure, host-independent helper so both
branches can be unit tested, including a regression test asserting that a host
without RDTSC reports the Stopwatch frequency rather than the hardware TSC.
2026-07-16 00:11:46 +03:00
Eyoatam Wubeshet 5c8ace82c0 Add missing French translations for Environment and About sections (#206)
Fills in the untranslated keys under Options.Env.* and About.* so the
French locale matches en.json. Tried to keep the wording consistent with
the rest of the file.

Co-authored-by: Rick21-bit <Rick21-bit@users.noreply.github.com>
2026-07-16 00:09:16 +03:00
can 341c2a0cb6 Fix sceRtcConvertLocalTimeToUtc failing on non-UTC hosts (#210)
The guest local tick was decoded into a DateTimeKind.Utc DateTime and then
passed to TimeZoneInfo.ConvertTimeToUtc together with TimeZoneInfo.Local.
That overload throws ArgumentException when a Utc-kind value is paired with a
source zone other than UTC, so on any machine whose local zone is not UTC the
export caught the exception and always returned INVALID_ARGUMENT.

Re-tag the decoded value as DateTimeKind.Unspecified so it is interpreted as
local wall-clock time and converted correctly. The reverse direction
(sceRtcConvertUtcToLocalTime) was already correct because ConvertTimeFromUtc
accepts a Utc-kind input.

Add a RtcExports unit-test suite covering the tick/calendar conversions, DOS
time packing, Win32 file time, leap-year and validation error codes, and a
regression test that round-trips a UTC tick through local time and back.
2026-07-16 00:04:35 +03:00
Gutemberg Ribeiro 320dbcacba [SourceGenerators] Compile-time SysAbi export registry, analyzers, and build-generated aerolib.bin (#204)
* [SourceGenerators] Add the SysAbi export generator and analyzers (phase 0)

New SharpEmu.SourceGenerators Roslyn component, complete and tested but
consumed by nothing yet — the emulator projects adopt it in the
following commits.

Ps5Nid ports the PS NID derivation (base64 of the byte-reversed first
eight SHA1 bytes of name + fixed suffix) from
scripts/generate_aerolib_binary.py to C#, so what has always been a
manual, out-of-band computation becomes a compile-time capability.

SysAbiExportGenerator emits a per-assembly SysAbiExportRegistry whose
CreateExports(Generation) reproduces ModuleManager's reflection scan
exactly — same generation inheritance and filtering, same method-name
fallback, same libKernel default — with attribute-omitted NIDs derived
algorithmically (equivalent to the runtime catalog lookup, which is
built from the same computation). Parameterless handlers are adapted to
the SysAbiFunction shape; invalid declarations are skipped here because
the analyzer rejects them as build errors, so nothing drops silently.

SysAbiExportAnalyzer turns the runtime failure modes into diagnostics:
SHEM001 duplicate NID (across declared and derived forms), SHEM002
malformed NID, SHEM003 uncallable handler signature, SHEM004 NID
contradicting its export name (the class of drift previously fixed by
hand), SHEM005 unresolvable export, SHEM006 export name unknown to
ps5_names.txt when the catalog is wired as an AdditionalFile, SHEM007
handler not reachable by generated code.

The self-contained test suite drives both in-process against the real
SharpEmu.HLE metadata: known catalog NID pairs pin the algorithm, the
generated registry must itself compile, and each diagnostic has a
triggering fixture. Fittingly, the NID pinning test caught a wrong
pair in its own first draft — the exact mistake SHEM004 exists to stop.

* [SourceGenerators] Adopt the generated export registry in the emulator (phase 1)

SharpEmu.Libs consumes the generator and analyzers, with
scripts/ps5_names.txt wired as the AdditionalFile catalog. The runtime
now registers exports from the compile-time SysAbiExportRegistry
instead of the boot-time reflection scan; RegisterFromAssembly is
retained solely as the arbiter for a parity test that pins the two
tables identical — same NIDs, names, libraries, targets, and handler
methods — across Gen4, Gen5, and combined registration.

First contact between the analyzer and all 715 existing exports
surfaced real drift the old offline checker structurally missed
(scripts/check_sysabi_aerolib.py skipped any NID absent from
aerolib.bin): three exports whose friendly names collide with real
catalog symbols of different NIDs, now suppressed at-site with reasons
pending AGC API confirmation, alongside the established synthetic
Unknown* labels for uncatalogued NIDs, which prompted a rule
refinement — SHEM004 only hard-errors when the export name is a real
catalog symbol, since synthetic labels cannot be validated by hashing
and the NID is authoritative for them. The two allowlisted mismatches
in the python checker no longer trigger anything, and the checker is
deleted: the analyzer subsumes it with the semantic model instead of
regex, and validates every declared pair rather than only
catalog-known NIDs.

* [SourceGenerators] Generate aerolib.bin at build time from ps5_names.txt

The runtime NID -> name catalog is derived data and no longer lives in
the repository: a Framework-only MSBuild task (GenerateAerolibBinaryTask,
sharing the same Ps5Nid implementation the analyzers use) builds it
into the intermediate directory from scripts/ps5_names.txt — now the
single source of truth — and SharpEmu.HLE embeds it from there. The
output is byte-identical to the previously committed binary, verified
with cmp against git history; a new test pins that the embedded catalog
loads and resolves a known symbol both directions.

scripts/generate_aerolib_binary.py is deleted (its algorithm lives in
Ps5Nid, its invocation in the build); the REUSE annotation for the
binary goes with it. MSBuild's Inputs/Outputs check means the ~154k NID
hashes only recompute when the names file actually changes. The task
implements ITask against Microsoft.Build.Framework directly, keeping
the vulnerable-flagged Utilities.Core package out and the analyzer
project's file-IO ban suppressed only inside the task itself.

* [SourceGenerators] Emit typed-signature register thunks (phase 2)

[SysAbiExport] handlers can now be written with real signatures — a
CpuContext followed by up to six int/uint/long/ulong parameters — and
the generator emits the SysV unmarshalling thunk, mapping parameters
positionally to RDI/RSI/RDX/RCX/R8/R9 with the same unchecked-cast
idiom hand-written handlers use. SHEM003 accepts the new shape and
rejects register overflow and non-register-representable types. Both
shapes coexist, so migration is per-handler; sceKernelPollSema,
sceKernelSignalSema, and sceKernelCancelSema migrate as the
demonstration (the last showing raw ulong guest-address passthrough).

The reflection scan cannot represent typed handlers, so it retires
here: RegisterFromAssembly, its signature validation, and
ResolveExportInfo are deleted, and the parity test that pinned the
generated registry to the scan is replaced by content-invariant tests
(duplicate-free, full 715-export surface, catalog identity). Deleting
the scan surfaced a phase-1 latent regression — the pre-JIT warm sweep
enumerated only reflection-scanned assemblies, so the generated
registration path warmed nothing and re-exposed the guest-thread
fail-fast risk; the warm set is now derived from the registered
handler delegates themselves.

* [SourceGenerators] Marshal guest strings declaratively with [GuestCString] (phase 3)

A string parameter on a typed [SysAbiExport] handler, annotated
[GuestCString(maxLength)], now makes the generated thunk read the
null-terminated UTF-8 string from the argument register's guest
address before the handler runs, returning
ORBIS_GEN2_ERROR_MEMORY_FAULT to the guest when the read fails —
the exact prologue nearly every string-taking handler writes by hand.
The attribute lives in SharpEmu.HLE next to SysAbiExportAttribute;
SHEM008 rejects misuse (non-string parameter, non-positive MaxLength)
while a bare string parameter stays a SHEM003 signature error.

_open, open, and sceKernelOpen migrate as the demonstration; they were
chosen because their hand-written prologue faulted on a null pointer
the same way the thunk does (handlers that return INVALID_ARGUMENT for
null pointers, like sceKernelCreateSema, keep the raw shape so guest-
visible semantics stay untouched).

* [SourceGenerators] Apply review findings across the branch

Behavior: the open/_open/sceKernelOpen [GuestCString] demo migration is
reverted — the local compat reader falls back to host memory for paths
in loader-mapped regions that ctx.Memory cannot see, so the generated
thunk would have turned recoverable reads into MEMORY_FAULT. The
marshalling infrastructure stays, proven by generator/analyzer tests;
production migration waits for a handler whose semantics the thunk
reproduces exactly. A comment on the handler records why.

Build robustness: the aerolib target is skipped for design-time builds
(the IDE resolves project references without compiling them, so on a
fresh clone the task assembly does not exist yet), and the task/names
paths are centralized in properties. The generator now emits no
registry for export-free assemblies, so referencing the analyzer can
never mint a colliding SharpEmu.Generated type.

Cleanup and perf: the pragma-suppression sites left mis-indented by the
phase-1 relocation are reformatted and the restores moved after the
method body; the dead ExportsForTesting hook and its InternalsVisibleTo
are deleted; the aerolib task reuses one SHA1 instance across ~150k
names; the analyzer caches the parsed catalog per file snapshot instead
of re-parsing 150k lines every compilation start, shares the attribute
name constant with the generator, and computes the catalog-membership
check once.

* [CI] Run the test suites in the build workflow

The workflow compiled the test projects (they are in SharpEmu.slnx) but
never executed them. A solution-level dotnet test now runs between
build and publish, so any test failure fails the build — including the
AerolibCatalogTests/SysAbiRegistryTests that guard the build-generated
aerolib.bin and the generated export registry. Generation failures of
aerolib.bin itself already fail the build step: the MSBuild task logs
an error event and returns false, and a missing task assembly or
missing embedded output are hard MSBuild errors. The NuGet cache key
now also tracks the test projects' lock files.

* [SourceGenerators] Address review feedback

Multi-diagnostic analyzer tests no longer assume a stable diagnostic
order (analyzer execution is concurrent), and the aerolib task logs
the full exception instead of only its message so build failures keep
the type and stack trace.

* [SourceGenerators] Address second review round

Symbol-name comparisons in the shape rules and analyzer now pin an
explicit SymbolDisplayFormat.FullyQualifiedFormat instead of relying on
the display-format default, and the aerolib task fails loudly on a
symbol name that would overflow the format's ushort length prefix
instead of silently truncating it, with null-safe output-directory
handling made explicit.

* [SourceGenerators] Embed aerolib.bin via a target so design-time builds never reference it

The static EmbeddedResource item referenced the generated file even in
design-time builds, where the generation target is skipped — on a fresh
clone the IDE would try to embed a file that never existed. The item is
now created inside an EmbedAerolibBinary target gated on
DesignTimeBuild, separate from the generation target so an up-to-date
skip of GenerateAerolibBinary cannot drop the item with the rest of its
body, and hooked before AssignTargetPaths since dynamic resource items
added later miss the resource pipeline.

Verified fresh build, incremental rebuild (embedded catalog test both
times), and a simulated design-time compile with no artifacts present.

* [SourceGenerators] Regenerate test lock file after rebase onto main

Rebase fallout: main's package graph shifted under #200, so the
SourceGenerators.Tests lock file is re-evaluated to keep --locked-mode
restore green at the branch tip.

* [Build] Drop NuGet lock files; rely on central package management

Central package management was already in effect (ManagePackageVersionsCentrally
with all versions in Directory.Packages.props and no inline PackageReference
versions), so the per-project packages.lock.json files and the lock-mode
workflow only added maintenance overhead. This removes all eleven lock files,
drops RestorePackagesWithLockFile so restore no longer regenerates them, and
takes --locked-mode off the CI restore steps (re-keying the NuGet cache on the
central props files). Package versions remain centrally pinned in
Directory.Packages.props.
2026-07-16 00:00:32 +03:00
Gutemberg Ribeiro 30fdd8d6ed [Gpu] Backend-neutral shader compiler and guest-GPU renderer seam (#200)
* [ShaderCompiler] Extract the backend-neutral shader compiler project

Move the Gen5 (gfx10) microcode decoder, the scalar evaluator, the
shader IR, and the metadata reader out of SharpEmu.Libs/Agc into a new
SharpEmu.ShaderCompiler project — the half of shader compilation every
codegen backend (SPIR-V today; MSL and DXIL later) consumes. Types go
public: they are the contract now. Nothing in the project may depend on
a host graphics API; the SPIR-V-specific artifact types
(Gen5SpirvShader, Gen5SpirvStage) stay beside the emitter in Libs.

Three couplings surfaced by the move, each resolved at the right depth:
GuestDrawKind was defined inside VulkanVideoPresenter despite being a
guest-domain, decoder-produced concept — it moves to the shared project;
the evaluator's one HLE dependency (the tracked-libc-heap read
fallback) becomes an injectable hook that a Libs module initializer
installs before any caller can reach the evaluator; and the inline-
constant table is promoted to a shared Gen5InlineConstants so backends
cannot drift on constant semantics (the SPIR-V translator now delegates
to it).

The ShaderDump tool drops its reflection over the moved types in favor
of direct typed calls; only the SPIR-V emitter, still internal to Libs
until it moves to its own backend project, is reached via reflection.
Verified by a clean solution build, the existing test suite, and a full
ShaderDump conformance run.

* [ShaderCompiler] Move the SPIR-V emitter into SharpEmu.ShaderCompiler.Vulkan

Gen5SpirvTranslator (with its ALU partial), SpirvModuleBuilder,
SpirvFixedShaders, and the Gen5SpirvShader/Gen5SpirvStage artifact types
move whole from SharpEmu.Libs/Agc into the first per-backend codegen
project. Notably it needs no Vulkan bindings reference: emitters
produce bytes from the shared IR; renderers own graphics APIs. Types go
public as the backend's contract; AgcExports and the presenter consume
them exactly as before.

The ShaderDump tool drops its last reflection: with both halves of the
pipeline public it drives decode and all three emit entry points with
direct typed calls, retiring the PadWithDefaults invoke shim — and it
no longer references SharpEmu.Libs at all, making the conformance tool
emulator-independent by design. Verified by a clean solution build, the
test suite, a full ShaderDump conformance run, and a locked-mode
restore under the pinned SDK.

* [Gpu] Extract the guest-GPU backend seam (IGuestGpuBackend)

The AGC/VideoOut/SystemService export layers now reach the renderer
through IGuestGpuBackend via GuestGpu.Current (mirroring HostPlatform),
instead of calling VulkanVideoPresenter statics. The Vulkan backend is
a thin adapter over the existing presenter, so the extraction stays
mechanical; only the adapter and the presenter itself reference the
presenter now.

The types crossing the seam move to Gpu/GuestGpuTypes.cs and drop their
Vulkan prefixes, which an audit showed were misnomers: every field is a
neutral primitive or a raw guest value (guest addresses, format and
number-type codes, CB_BLEND register bitfields, verbatim sampler
descriptor dwords). The one genuine Vulkan value in the old surface —
the Silk.NET Format inside VulkanRenderTargetFormat, which callers
never read — stops crossing: TryDecodeRenderTargetFormat is replaced at
the seam by TryGetRenderTargetOutputKind, which surfaces only the
Gen5PixelOutputKind callers actually consume, keeping native formats a
backend-internal concern. ToVulkanSampler in AgcExports is renamed
ToGuestSampler to match what it always produced.

Seam rules are documented on the interface: no host-API value crosses,
and submission stays coarse-grained with synchronization internal to
backends. Interim exception, resolved next: shader parameters are still
SPIR-V blobs.

* [Gpu] Move shader compilation behind the guest-GPU backend

The seam's interim exception is gone: AgcExports no longer calls
Gen5SpirvTranslator or handles SPIR-V bytes. IGuestGpuBackend gains the
three TryCompile entry points, which take the backend-neutral
(Gen5ShaderState, Gen5ShaderEvaluation) contract plus the flat
per-role resource-slot bases a multi-stage draw needs, and return
opaque IGuestCompiledShader handles that only the producing backend can
submit — the Vulkan backend wraps its SPIR-V in
VulkanCompiledGuestShader and rejects foreign handles loudly. Draw and
dispatch submissions take handles instead of byte arrays; the shader
caches in AgcExports store handles.

IGuestCompiledShader.Payload exposes the backend-defined compiled bytes
for exactly two callers: the diagnostics dump and the size trace —
documented as never-interpret. The unused _pixelSpirvCache is deleted.
With this, a Metal or DX12 backend plugs in by implementing
IGuestGpuBackend with its own codegen; nothing in the export layers
knows which shader format exists.

Verified by a clean solution build, the test suite, and a full
ShaderDump conformance run under the pinned SDK.

* [Gpu] Fix rename collateral from the seam extraction

Address review findings: a doc comment picked up the mechanical
VulkanVideoPresenter -> GuestGpu.Current rewrite and ended up naming
members that do not exist on the interface, and CreateVulkanIndexBuffer
kept its Vulkan prefix while every sibling factory was de-Vulkanized —
it produces the neutral GuestIndexBuffer, so it is CreateGuestIndexBuffer.

* [Gpu] Label diagnostics dumps with the backend's payload extension

Address the review's altitude finding on DumpSpirv: the dump helper's
IR-disassembly half is backend-neutral and stays put, but writing the
opaque payload to a hardcoded .spv interpreted bytes the seam says
never to interpret. IGuestCompiledShader now declares its payload's
file extension, and the renamed DumpCompiledShader takes the handle and
writes honestly-labeled dumps whichever backend produced them.

* [Gpu] Make the shader-cache hit path allocation-free and lock-free

Every translated draw built its cache key with a LINQ Select feeding
string.Join plus one interpolated string per render target — steady
per-draw allocation whether or not the shaders were already cached. The
output layout is now packed exactly into a ulong (guest slot in 6 bits
+ output kind in 2 bits per target, host locations being the byte
positions, target count in the key beside it), and the
Gen5PixelOutputBinding array is only materialized on a cache miss,
where compilation dwarfs it.

The graphics/compute shader caches switch from Dictionary guarded by
_submitTraceGate to ConcurrentDictionary, making the per-draw and
per-dispatch hit paths lock-free and decoupling them from the tracing
gate they coincidentally shared. And the seam-shaped render-target list
is built once when a translated draw is created instead of a
Select/ToArray per submission of a cached draw.

* [Gpu] Replace LINQ with explicit loops in code this branch introduced

Project rule going forward: no LINQ — it allocates enumerators,
closures, and delegates, and this codebase is GC-pause-sensitive. The
pixel-output and guest-render-target array builds and the ShaderDump
store-PC collection become plain loops; pre-existing LINQ elsewhere is
left for changes that already touch those lines.

* [ShaderCompiler] Suppress CA2255 on the evaluator hook installer

The analyzer coverage that arrived with the rebase flags
ModuleInitializer in library code; this is the rule's intended advanced
scenario — the hook must be installed before any code path can reach
the evaluator, and every such path enters through this assembly — so
suppress with that justification rather than weaken the guarantee to a
static constructor's lazier timing.

* [Gpu] Resolve rebase artifacts onto main

Dedupe the System.Collections.Concurrent using in AgcExports that the
rebase merge duplicated (main and this branch each added it), and
regenerate the lock files for the new shader-compiler projects and
SharpEmu.Libs against main's current package graph so --locked-mode
restore matches at the branch tip.

* [CI] Comment per-platform build artifact links on PRs

Adds a workflow_run workflow that, after "Build and Release" finishes a
pull-request build, posts (and keeps updated in place) a single PR
comment linking the Windows, Linux, and macOS artifacts from that run.

It runs via workflow_run rather than in the build workflow because PRs
from forks build with a read-only token that cannot comment; the
follow-on run executes in the base-repo context with write access and
without checking out fork code. GitHub only triggers workflow_run from
the default branch, so this takes effect once merged to main.
2026-07-15 11:11:24 -06:00
Berk c69ac6ddab [CI] Update for cross-platform support (#209)
* [CI] Update for cross-platform support

* [CI] fix check
2026-07-15 16:07:36 +03:00
kuba fa2616d224 Linux and macOS support (#47)
* [macos/linux] Cross-platform host memory, TLS, and ABI layer for POSIX

Introduces the foundation for running SharpEmu on macOS (osx-x64 under
Rosetta 2) and Linux (linux-x64). The CPU backend executes guest x86-64
code natively, so these targets run the whole process as x86-64; this
commit replaces the Windows-only host primitives with platform-dispatched
equivalents so the guest boots and services HLE calls off Windows.

Memory (HostMemory.cs, new): a Win32-semantics facade over
mmap/mprotect/munmap with a shadow region table answering VirtualQuery.
PhysicalVirtualMemory, DirectExecutionBackend, StubManager, and the two
Kernel*CompatExports now go through it instead of kernel32 P/Invokes.
Exact-address requests use MAP_FIXED_NOREPLACE (Linux) / guarded
MAP_FIXED (macOS) so they match Win32 "map there or fail" semantics.

TLS + host helpers (PosixHostStubs.cs, new): pthread-backed TLS and
Win64-ABI-compatible stubs for the kernel32 helpers the backend embeds
into emitted x86-64 code (TlsGetValue, QueryPerformanceCounter,
SwitchToThread, Sleep). A Win64->SysV thunk wraps managed callbacks,
since .NET on POSIX compiles them for the SysV ABI while the emitted
call sites use Win64.

Guest address layout: the 0x7FFx window is Windows-only (dyld shared
cache / Rosetta runtime live there on POSIX), so stack/TLS/stub regions
relocate to 0x6FFx off Windows.

Vectored exception handling is gated off on POSIX for now (guest faults
are not yet recovered) — the signal-based bridge is the next step. Also
adds osx-x64 to the RID list and a Docker-based Linux smoke-test script.

Status: on both macOS (Rosetta) and Linux (amd64), the guest now boots,
runs native x86-64 code, and dispatches HLE imports. macOS stops at a
Rosetta translation-cache issue; Linux runs ~252 imports through C++
static-init before hitting the missing fault handler (SIGSEGV).

* [posix] Bridge the vectored exception handler to sigaction(SIGSEGV/SIGBUS/SIGILL)

Guest faults on macOS/Linux previously terminated the process because the
recovery logic in DirectExecutionBackend.Exceptions.cs was Windows-only.
This adds a POSIX front-end that reuses the existing handler bodies:

- DirectExecutionBackend.PosixSignals.cs installs SA_SIGINFO handlers via
  an [UnmanagedCallersOnly] entry, rebuilds the Win64 EXCEPTION_POINTERS /
  CONTEXT view from the platform mcontext (Darwin __ss thread state via
  the mcontext pointer at ucontext+48, Linux glibc gregs at ucontext+40 --
  offsets verified against the headers on both platforms), runs the same
  chain as the VEH path (TryRecoverUnresolvedSentinel trap-sentinel
  recovery, TryHandleLazyCommittedPage demand paging, VectoredHandler
  diagnostics incl. FS/GS TLS-fault detection), and writes register
  changes back into the mcontext so sigreturn resumes the repaired guest.
  Unrecovered faults chain to the previously installed handler so the
  .NET runtime keeps mapping its own faults to managed exceptions.

- The whole recovery path is warmed up with fabricated inputs before the
  handlers are installed. This is required under Rosetta 2: the signal
  trampoline cannot enter x86 code that has never been executed (and so
  never translated) -- a cold handler is silently never invoked and the
  faulting instruction retries forever (reproduced and verified in an
  isolated .NET test under Rosetta for Linux). It also keeps first-fault
  JIT work out of the signal frame.

- Handlers run without SA_ONSTACK: the runtime's alternate stacks are too
  small for the diagnostic path, while guest (2MB) and host thread stacks
  match where Windows dispatches exceptions anyway.

- The raw reads in the shared fault diagnostics (stack qwords, RBP walk,
  code bytes at RIP) now probe the region table on POSIX before touching
  memory, since a nested SIGSEGV inside the handler would kill the
  process before diagnostics finish. Windows keeps its try/catch reads.

- Escape hatches: SHARPEMU_DISABLE_POSIX_SIGNALS=1 skips installation,
  SHARPEMU_DISABLE_RAW_HANDLER=1 disables sentinel recovery (parity with
  Windows), SHARPEMU_LOG_POSIX_SIGNALS=1 traces every delivery (first 16
  and every 1024th are always traced).

Verified with the test game: Linux (amd64 container) previously died with
SIGSEGV right after import #252; it now recovers/diagnoses signals and the
run proceeds to the real next blocker, an unpatched negative-offset guest
TLS read (fault at TLS base - 0x1708), which gets the full NATIVE
EXCEPTION dump before terminating. macOS is unchanged: the bridge installs
and the game still stops at the known Rosetta translation-cache error at
import 12, which is the next work item.

* [posix] Fix guest memory layout faults: TLS prefix, exact mmap, map search base

Three fixes that take the test game from dying during libc init to running
its full main loop on macOS and Linux:

- Static TLS blocks live below the TCB (FreeBSD amd64 variant II) and
  libc.prx reaches past -0x1700, but only a 4KB prefix was mapped below
  the TLS base. The prefix is now 64KB on POSIX (Windows keeps 4KB); the
  fault was a read at TLS base - 0x1708 during libc init.

- HostMemory exact allocation on macOS used MAP_FIXED, which silently
  maps over untracked host memory. The direct-memory allocator's address
  scan walked into the .NET runtime's JIT heap and replaced live code,
  which under Rosetta 2 surfaced as "no code fragment associated with
  the given arm pc". Exact placement now passes the address as a hint
  and fails on relocation, like MAP_FIXED_NOREPLACE does on Linux.

- sceKernelMapDirectMemory/MapFlexibleMemory searched for free space
  starting at 4GB, which is the Mach-O image base on macOS. The default
  search base is 0x20_0000_0000 on POSIX, and TryAllocateAtOrAbove now
  asks the kernel for a placement instead of page-stepping through host-
  owned address space (Rosetta ignores mmap hints for whole VA windows),
  over-allocating when the caller needs more than page alignment.

Windows behavior is unchanged; all divergences are platform-guarded.

* [macos] Video presenter on the main thread, MoltenVK support, window keyboard input

Gets the test game from a headless loop to a playable window on macOS:

- AppKit traps with SIGILL ("NSUpdateCycleInitialize() is called off the
  main thread") when GLFW runs on a worker thread. The CLI now moves
  emulation onto a worker thread on macOS and parks the real main thread
  in HostMainThread.Pump(); the presenter posts its whole window loop
  there instead of spawning a thread, and a shutdown handler asks the
  render loop to close the window so the pump unwinds on guest exit.

- MoltenVK: enable VK_KHR_portability_enumeration (+ the portability
  instance flag) and VK_KHR_portability_subset when advertised, and gate
  robustBufferAccess2 on the device actually supporting it (Metal does
  not; the old code keyed it off robustImageAccess2 and vkCreateDevice
  failed with ErrorFeatureNotPresent).

- Input: pad exports polled user32 GetAsyncKeyState, so POSIX hosts threw
  DllNotFoundException per scePadReadState call. The presenter now
  attaches the window's keyboard via Silk.NET.Input into HostWindowInput,
  and the pad exports map the existing VK-code layout onto it off
  Windows. Headless hosts (Linux containers) report a disconnected
  keyboard and fall back to neutral pad data silently.

GLFW needs an x86-64 Vulkan loader under Rosetta: place a universal
libMoltenVK.dylib next to SharpEmu named libvulkan.1.dylib (Homebrew's
arm64-only copy cannot load into the x86-64 process) and export
DYLD_LIBRARY_PATH to that directory.

Verified: Dreaming Sarah boots to a MoltenVK-backed 2560x1440 window on
macOS (Apple M4, Rosetta 2), renders the intro, title, and menus, and
keyboard input drives it into gameplay. Linux (amd64 container) runs the
same build headless through millions of imports with no faults. Windows
paths unchanged; arm64 and x64 builds clean.

* [posix] CoreAudio playback, self-contained MoltenVK loading, input/log polish

- Audio: sceAudioOut ports now play through an AudioQueue backend on macOS
  (stereo PCM16 with the same 32KB backpressure pacing as the WinMM path).
  The WinMM port and the new CoreAudio port share an IHostAudioPort
  interface and sample converter; hosts without a backend (Linux
  containers) keep the silent fallback.

- MoltenVK: GLFW resolves Vulkan with dlopen("libvulkan.1.dylib"), which
  cannot see the app-local universal MoltenVK build, so the presenter now
  feeds vkGetInstanceProcAddr straight into glfwInitVulkanLoader (GLFW
  3.4) before creating the window. No DYLD_LIBRARY_PATH needed; the CLI
  also preloads the dylib for Silk.NET and prints setup hints when it is
  missing. scripts/fetch-macos-moltenvk.sh stages the official universal
  dylib next to a build.

- The virtual-range allocator's failure trace now names the address and
  length instead of "AllocateAt invocation threw".

Investigated and documented (not port defects): the savedata transaction
failure is identical on Linux and macOS (HLE argument-register mapping for
sceSaveDataCreateTransactionResource), and the in-game tile speckling has
no platform-specific code in its path - the one macOS-only delta is that
MoltenVK lacks robustBufferAccess2, so out-of-bounds shader reads return
garbage instead of zeros.

Verified on macOS: window, audio backend, and keyboard input all come up
with zero environment configuration; the game runs to gameplay. Linux
headless run unchanged (silent audio, no faults). Windows paths untouched;
arm64 and x64 builds clean.

* [cpu] Preserve guest registers and flags across patched TLS accesses

The TLS patch handler replaces guest `mov reg, fs:[...]` instructions,
which preserve every other register and the flags - but the handler
loaded the TLS index into ecx and called TlsGetValue (Win64: clobbers
rcx/rdx/r8-r11) with `sub/add rsp` trashing the arithmetic flags. Guest
code that keeps live values or comparison results across a TLS access
computed garbage deterministically. The handler now saves rcx, rdx,
r8-r11, and the flags around the call, keeping the same inner stack
alignment. This applies to the load patches and both store-helper stubs,
on every platform.

Also in this change, from the rendering-artifact investigation:

- The present blit picks linear filtering for any fractional scale
  (nearest only for integer upscales): a 3840x2160 guest frame blitted
  into a 2560x1440 swapchain with nearest silently dropped every third
  row/column.
- ClampViewport no longer trims the guest viewport rectangle to the
  render target; trimming changed the guest's scale/offset and skewed
  texel addressing. Vulkan permits viewports beyond the framebuffer
  (the scissor confines rendering), so only spec bounds are enforced.
- Env-gated diagnostics grown during the investigation: guest texture
  dumps (SHARPEMU_TEXTURE_DUMP_DIR), aliased guest-image readback dumps
  (SHARPEMU_TRACE_GUEST_IMAGES=alias), small-render-target write movies
  (SHARPEMU_TRACE_GUEST_WRITES=small), unattended input injection
  (SHARPEMU_AUTO_CROSS=secs,...), viewport nudging
  (SHARPEMU_VIEWPORT_EPSILON), chunked-draw toggle
  (SHARPEMU_DISABLE_CHUNKED_DRAWS), and rect-list/draw vertex traces.

Known remaining issue (root cause narrowed, not yet fixed): the game's
terrain texture pages are corrupted in guest memory before any GPU work
- the mound's solid-fill 32x32 tiles decode to fully transparent texels
and the grass page has deterministic gaps, byte-identical across runs.
Ruled out: memcpy/memmove/memset/realloc HLE semantics, sampler wrap
modes, texel-boundary rounding, chunked draws, viewport handling. Next
step is auditing the Chowdren asset decode path (custom compressed
images) against the emulator's import surface.

* [linux] ALSA playback backend for sceAudioOut

sceAudioOut ports on Linux now play through libasound instead of the
silent fallback. The PCM device opens in blocking mode with ~170ms of
device buffer (the time-equivalent of the 32KB queue the WinMM and
CoreAudio ports keep), so snd_pcm_writei provides the same backpressure
pacing without a managed queue. Underruns and suspend/resume go through
snd_pcm_recover with one retry per submit; anything else drops the
buffer rather than stalling the guest.

The "default" device routes through PulseAudio/PipeWire on desktops
and straight to hardware on bare ALSA; SHARPEMU_ALSA_DEVICE overrides
it (the null device makes the path testable in containers). A missing
libasound or device fails port creation and lands in the existing
silent fallback.

Verified in an amd64 container: the test game opens the port
(backend=alsa, 48kHz stereo float32) and streams sceAudioOutOutput
through the null device for a full run; without a usable device the
port logs a warning and falls back to silent. Playback on real Linux
audio hardware has not been tested.

* [fixes] Address review feedback: commit bounds, CoreAudio shutdown, dump errors

- HostMemory: a MEM_COMMIT that runs past its reservation now fails like
  Win32 instead of committing a prefix and reporting success. All current
  callers already clamp their ranges to the region, so this only guards
  future callers.

- CoreAudioPort: Dispose wakes a submitter waiting on backpressure and
  the wait treats ObjectDisposedException as a timed-out wait, so closing
  a port during playback can no longer throw. A failed AudioQueueStart
  tears the queue down and fails fast instead of leaving an undrainable
  queue that stalls every later submit on its timeout.

- AgcExports: texture dumping catches all write failures (bad path,
  permissions), logging a warning instead of crashing when
  SHARPEMU_TEXTURE_DUMP_DIR points somewhere unusable.

Verified with the Linux container run: game boots and streams audio with
the stricter commit check, and a dump dir under /proc produces warnings
instead of taking the process down.

* [ci] Build linux-x64 and osx-x64 archives

Adds a build-posix matrix job (ubuntu-latest / macos-latest) mirroring
the Windows build: locked restore, Release build, self-contained CLI
publish, and a tar.gz artifact per RID (tar keeps the executable bit).
The macOS archive also stages the universal MoltenVK dylib via
scripts/fetch-macos-moltenvk.sh so the artifact runs without any manual
Vulkan setup. The release job still only ships the Windows archive.

* [cli] Keep POSIX glfw natives outside the single-file bundle

The KeepGlfwOutsideSingleFile target only matched filenames starting
with 'glfw', which covers Windows (glfw3.dll) but not libglfw.3.dylib /
libglfw.so.3. Those got embedded into the single-file bundle, and
Silk.NET's library loader does not probe the bundle extraction
directory, so a published build died with "Couldn't find a suitable
window platform" (and the glfwInitVulkanLoader wiring, which loads the
library from AppContext.BaseDirectory, could not run either). Keeping
the POSIX names loose next to the executable fixes both, the same way
the Windows build already handled it.

Found by running the CI-built osx-x64 archive: video failed while local
loose-file builds worked. With the fix the published single-file build
opens the MoltenVK window, wires the loader, and reaches gameplay.

* [ci] Publish linux-x64 and osx-x64 release archives

The build-posix artifacts now ship as per-RID GitHub releases on main
pushes and manual dispatches, tagged the same way as the win64 ones
(<rid>-<ref>-<sha>). Archives stay tar.gz so the executable bit
survives extraction.

* [cli] Fail early on non-x86-64 host processes

The CPU backend executes guest x86-64 code natively, so the process
must be x86-64 (win-x64/linux-x64 on x64 hardware, osx-x64 under
Rosetta 2 on Apple Silicon). An arm64 process previously failed deep
inside emulation startup, indistinguishable from MoltenVK, signal
handler, or guest memory problems. CLI mode now checks the process
architecture up front and exits with a message naming the supported
execution model (and the Rosetta install command on macOS). The
GUI-only path stays usable on arm64.

* [video] Log the selected Vulkan device name and API version

The presenter never named the GPU it picked, so a 'no video' report
could not be told apart from a real windowing failure without guessing.
It now logs the device name, type, and API version right after
selection. A software rasterizer (llvmpipe/lavapipe/SwiftShader) shows
up here and typically lacks the device features the translated shaders
need, which is the likely cause when a window opens and presents frames
but nothing draws.

* [video] Steer GLFW to XWayland on Wayland sessions

GLFW's native Wayland backend does not reliably map the Vulkan window
with some drivers (NVIDIA in particular): frames present but the window
never becomes visible, so the game runs with audio and no picture. A
report on an RTX 5080 showed exactly this — all device features present,
frames presenting, but the log had 'libdecor-gtk.so failed to init' and
a 1.4x-scaled window, both Wayland tells.

On a Wayland session that also exposes an X server (DISPLAY set), the
presenter now clears WAYLAND_DISPLAY for its own process before GLFW
initializes, so GLFW selects its dependable X11/XWayland backend.
SHARPEMU_ENABLE_WAYLAND=1 opts back into native Wayland. Headless
(no DISPLAY) and non-Linux hosts are unaffected.

* [video] Force GLFW X11 backend via the platform init hint, log the platform

The previous attempt cleared WAYLAND_DISPLAY to steer GLFW off Wayland,
but a reporter still hit the native-Wayland path (the Wayland-only
libdecor error persisted), so that env trick doesn't switch GLFW.

Use GLFW's supported mechanism instead: glfwInitHint(GLFW_PLATFORM,
GLFW_PLATFORM_X11) before GLFW initializes, called into the same libglfw
GLFW itself loads (the pattern InitializeMacVulkanLoader already uses).
Still gated on a Wayland session with an X server present (DISPLAY set)
so we never force X11 where XWayland can't catch it, and still
overridable with SHARPEMU_ENABLE_WAYLAND=1.

Also logs 'GLFW windowing platform in use: <backend>' after init via
glfwGetPlatform, so a 'no window' report shows X11 vs Wayland outright.
Verified on macOS: the readback correctly reports Cocoa and the
presenter is unaffected (the fix is a no-op off Linux).

* [video] Run the GLFW window on the main thread on Linux too

GLFW requires window creation and event processing on the process main
thread on every platform: initialization, window creation, and
glfwPollEvents are main-thread-only, and X11 in particular has a single
event queue that must be serviced there. A window created and polled on
another thread may never map — which is why the game ran (audio, imports,
even Vulkan present) with no visible window on Linux.

macOS already routed the window loop to the main-thread pump (AppKit
needs it); Windows is fine because it has a per-thread event queue. Linux
was the gap: it spawned a background thread for the presenter. Extend the
existing HostMainThread pattern to Linux — emulation runs on a worker,
the main thread pumps the window work the presenter posts.

Refs GLFW intro guide (thread-safety): init, window creation, and event
processing are restricted to the main thread.

Verified: macOS still boots to its window unchanged; the Linux headless
container runs to millions of imports with no deadlock or regression.
On-screen confirmation on a real Linux desktop is still pending, but this
is the documented root cause for a windowless-but-running Linux session.

* [posix] Skip Win32 native guest workers

* [vulkan] Synchronize offscreen targets before present

* [vulkan] Transition fresh textures from undefined layout

* [vulkan] Report swapchain pixels before source readback

* [vulkan] Emit requested guest image diagnostics

* [agc] Diagnose guest texture fallbacks

* [linux] Keep guest GPU mappings in low address space

* [video] Reduce diagnostic stalls and drain complete frames

* [memory] Harden packed GPU address handling

* [readme] Document Linux and macOS release support

* [posix] Integrate the host platform abstraction

* [posix] Restore guest thread address window

* [video] Run the performance HUD on POSIX hosts

The FPS/CPU/work HUD bailed out unless the host was Windows; only the
per-thread hottest-thread scan actually needs Windows APIs. Keep that
scan Windows-only (POSIX reports 'idle') and let the rest of the HUD
run everywhere — the title is already set from the render thread, which
owns the window on macOS and Linux.

* [posix] Implement native guest worker threads

Guest entry stubs must not run above CLR-managed frames on CLR-created
threads (see the NativeWorker preamble); the PR previously fell back to
the inline calli path on POSIX, which reproduced the documented
'attempted to call a UnmanagedCallersOnly method from managed code'
fail-fast (observed after Dreaming Sarah's menu select) and left the
runtime's suspension machinery walking guest frames.

Provide the missing POSIX half of the worker loop:
- PosixHostStubs grows Win64-convention WaitForSingleObject/SetEvent/
  ExitThread stubs backed by dispatch semaphores (macOS) / unnamed POSIX
  semaphores (Linux) plus pthread_exit, with EINTR retry in the wait.
- Worker events are creatable/signalable/waitable from managed code too,
  so NativeGuestExecutor.Run keeps its handshake (AutoResetEvent stays
  on Windows byte-for-byte).
- PosixHostThreading implements CreateNativeThread/WaitForThreadExit/
  CloseThreadHandle over pthreads (liveness probed with
  pthread_kill(0), then joined).
- RunPrologue/RunEpilogue are routed through the existing Win64->SysV
  thunks, so the emitted loop stays identical across platforms.

* [macos] Disable concurrent GC under Rosetta's write-watch hazard

Background GC's write-watch revisit (SoftwareWriteWatch::GetDirty ->
FlushProcessWriteBuffers) calls thread_get_register_pointer_values on
every thread; under Rosetta 2 that Mach call stalls indefinitely on
threads executing translated guest code. The background mark phase then
never finishes and every allocating or Monitor-taking thread wedges
behind it — observed as Dreaming Sarah freezing at the menu/loading
screen with FPS 0 in 5 of 7 runs, dispatcher/watchdog parked in
Monitor.Enter and all BGC threads waiting in t_join.

Non-concurrent GC never takes that path; a 5-minute soak now holds
22-31 fps in-game with zero stalls. Windows and Linux keep concurrent
GC.

* [diag] Periodic guest-thread snapshots with gate-owner tracking

SHARPEMU_PERIODIC_SNAPSHOT_SECONDS=N dumps the stall snapshot every N
seconds even while imports are progressing, for soft stalls where the
game stops advancing but threads keep spinning. The periodic dump never
touches the guest-thread gate (it must keep reporting when the gate is
what's wedged): it reads a lock-free owner record — every gate
acquisition now goes through LockGate(site), which notes site/thread —
and walks the thread table without the lock, tolerating torn reads.
SHARPEMU_PERIODIC_SNAPSHOT_FILE redirects the dump to a side file for
the case where the console itself is wedged (frozen log mirror was one
of the observed failure modes).

* [nuget] Add osx-x64 RID targets to lock files

* [cpu] Back off the guest join poll

TryJoinThread polled the host thread at a fixed 1ms; a game main thread
joining a long-lived worker (Dreaming Sarah parks there for the whole
session) burned ~5% of managed CPU in Join/Sleep syscalls. Ramp the
poll interval to 10ms once the join is clearly long-lived — exit
detection latency for long joins moves from ~1ms to at most 10ms, and
short-lived joins still resolve on the first 1ms polls.

* [nuget] Add linux-x64/win-x64 RID targets to lock files

* [posix] Keep guest stacks clear of the import-stub descent

The import-stub region descends from 0x7000_0000_0000 on the same 16MB
grid as the guest thread windows; moving stacks to 0x6FFF_E000_0000 put
them inside the stub region's 64-module descent range (floor
0x6FFF_C000_0000), silently consuming the top ~32 stack slots on hosts
with many loaded modules. Drop the POSIX stack base to 0x6FFF_A000_0000:
512MB below the stub floor, still 2.5GB above the TLS window. Windows
keeps 0x7FFF_E000_0000 (its bands are ~15TB apart).

* [pad] Read window gamepads on POSIX hosts

XInput and the DualSense hid reader are Windows-only, which left
macOS/Linux with keyboard input only. The presenter's Silk/GLFW input
context already enumerates gamepads on both platforms, so track their
state in HostWindowInput (event-driven on the window thread, snapshot
guarded like the key set) translated to ORBIS conventions: GLFW's Xbox
layout maps A/B/X/Y to Cross/Circle/Square/Triangle, sticks bias from
-1..1 to 0..255 with Y growing down, and triggers rescale from GLFW's
-1..1 resting-at--1 range with digital L2/R2 bits past 25%.

The merge into ReadHostInputState is gated to non-Windows so a physical
pad is never sampled twice through both a native reader and GLFW.
Hotplug is handled via ConnectionChanged; with no pad connected the
path is inert.

Untested against a physical controller (none attached to the dev host);
axis conventions follow the GLFW gamepad-mapping contract.

* [nuget] Refresh lock files after cross-RID restores

* [posix] Adopt the host audio/input seams from main

Main's #192 abstracted audio output and pad/keyboard input behind
IHostAudioOutput/IHostInput; re-express the POSIX backends behind them:

- CoreAudioPort/AlsaAudioPort move to Host/Posix as
  PosixCoreAudioStream/PosixAlsaAudioStream implementing
  IHostAudioStream. The seam converts to stereo PCM16 before Submit, so
  the ports' own conversion (and IHostAudioPort/AudioSampleConverter)
  is gone; queueing and backpressure are unchanged.
- PosixHostAudio selects CoreAudio (macOS) / ALSA (Linux) as the
  platform's IHostAudioOutput.
- PosixHostInput implements IHostInput over an
  IPosixWindowInputSource that HostWindowInput registers when the
  presenter attaches the window's GLFW input context: keyboard with
  virtual-key translation, the window gamepad snapshot (now in seam
  HostGamepadState/HostGamepadButtons terms), and keyboard-connected as
  the focus signal. Rumble/lightbar no-op (GLFW has no such API).
- PadExports drops its direct HostWindowInput gamepad merge — pads now
  flow through IHostInput.GetGamepadStates like every platform.
- PosixHostThreading.RequestTimerResolution is a documented no-op.

All three RIDs build; SharpEmu.Libs.Tests pass (26/26).

* [nuget] Regenerate GUI lock file for RID-less locked restore

Local cross-RID builds stamped a win-x64 runtimes section into
SharpEmu.GUI's lock file; the project declares no RuntimeIdentifiers,
so CI's 'dotnet restore --locked-mode' failed with NU1004 on every
platform. Regenerated via a plain solution restore (--force-evaluate),
matching what the workflow validates.
2026-07-15 15:36:20 +03:00
Pacuka 7b86a91dfa Add files via upload (#202)
Added Hungarian tranlation. -Pacuka
2026-07-15 14:39:42 +03:00
Gutemberg Ribeiro 72645cb373 [Host] Abstract audio output and pad/keyboard input behind the host platform seam (#192)
* [Host] Abstract audio output behind IHostAudioOutput

Add IHostAudioOutput (opens streams, names the backend for diagnostics)
and IHostAudioStream (submit interleaved stereo 16-bit PCM, Dispose) to
the host seam, with the winmm waveOut implementation moving whole into
Host/Windows/WindowsWaveOutAudio — same device open, queueing,
32 KB backpressure wait, and buffer lifetime as WinMmAudioPort had. The
DllImports become source-generated LibraryImports in the move, matching
the other Windows backends.

The guest-format conversion (mono/stereo/7.1, s16/float32 -> stereo
PCM16) is platform policy, not device code, so it stays in Libs as
AudioPcmConversion; AudioOutOutput converts into a pooled buffer and
submits the result through the stream. Open failures still degrade to
the silent paced port with the same warning, and the port log line now
takes its backend name from the platform instead of a hardcoded string.

* [Host] Abstract pad and keyboard input behind IHostInput

Add IHostInput to the host seam: gamepad state snapshots, rumble /
trigger-rumble / lightbar sinks, and the keyboard-fallback queries
(window focus, key state). Gamepad state crosses the seam as the new
unmanaged HostGamepadState with HostGamepadButtons flags — named after
the PlayStation layout the guest API exposes but with the seam's own
values, so SCE_PAD_BUTTON bits never leak into host backends and the
per-frame poll can stackalloc its snapshot buffer.

The DualSense raw-HID reader, the XInput reader, and the Win32 HID
interop move whole into Host/Windows (report parsing, hot-plug loops,
rumble/lightbar output reports, and log strings unchanged), translating
to the neutral flags instead of ORBIS bits and converting their
DllImports to source-generated LibraryImports. WindowsHostInput
composes them plus the user32 keyboard queries; rumble still fans out
to both readers, trigger rumble stays XInput-only, lightbar stays
DualSense-only.

PadExports keeps all policy: the keyboard mapping (now via named
OrbisPadButton constants instead of raw hex), the controller-beats-
keyboard-past-deadzone merge, and the new host->ORBIS button
translation. The GUI's source-linked reader copies re-point to the
moved files (it still cannot reference SharpEmu.HLE wholesale), which
requires AllowUnsafeBlocks for the generated marshalling stubs; its
navigation code switches to the neutral flags.

* [Host] Move the timer-resolution request behind IHostThreading

IHostThreading gains RequestTimerResolution (idempotent, best-effort
~1 ms timed-wait granularity; a no-op wherever the platform default is
already fine). The winmm timeBeginPeriod call, its once-only latch, and
both warning strings move from the Libs-level HostTimerResolution
helper into WindowsHostThreading as a source-generated LibraryImport;
the vblank pump requests it through the platform instead.

HostSystemInfo in SharpEmu.Logging keeps its direct user32/kernel32
imports deliberately: Logging sits below HLE in the dependency chain so
it cannot see the host seam, every path is already OS-gated with
fallbacks, and it only runs once for the diagnostics banner.
2026-07-15 13:24:53 +03:00
SamuelEzequias 2ad9836d13 Add Brazilian translation to Environment tab (#196) 2026-07-15 13:00:30 +03:00
Gutemberg Ribeiro 62e1775c5c [HLE] Remove steady-state allocations from the hot HLE paths (#190)
* [HLE] Stop allocating on the memcpy/memset and trace hot paths

memcpy/memmove no longer allocate a bounce buffer sized to the whole
copy (large copies previously landed on the LOH); they loop through a
single pooled 256 KB rental, copying high-to-low when the destination
overlaps above the source so memmove semantics survive the chunking.
memset reuses a shared zero chunk for the dominant zero-fill case and
rents/fills only min(length, 16K) bytes for non-zero values instead of
allocating and filling a fresh 16 KB array per call; the map-time
zero-fill loop shares the same zero chunk.

SHARPEMU_LOG_SEMA / SHARPEMU_LOG_VIDEOOUT are now read once into cached
bools and every TraceSemaphore/TraceVideoOut call site is guarded, so
trace messages are no longer interpolated (and the env var no longer
queried) on every semaphore op and every flip with tracing off. Trace
output when the flags are set is unchanged.

* [HLE] Remove per-frame allocations from the vblank/flip/equeue plumbing

The 60 Hz vblank pump no longer allocates per edge: PumpVblanks reuses a
pump-thread-only port list instead of a LINQ Where/ToArray, and
SignalVblank/SubmitFlip snapshot their event registrations into pooled
rentals instead of copying the List on every edge and every flip (the
snapshot must still be taken, since triggers run outside _stateGate and
a per-port reusable buffer would race the pump thread against a guest
thread's first-edge signal).

sceKernelWaitEqueue delivery rents the dequeue buffer from the pool
instead of allocating an array per wait, and event-queue wake keys are
formatted once per handle (cached in a ConcurrentDictionary, dropped on
queue delete) instead of building the string on every enqueue. The
semaphore wake key moves onto KernelSemaphoreState at creation, the
same pattern the pthread mutex state already uses, removing the
per-signal/per-wait formatting. SHARPEMU_LOG_EQUEUE is read once into a
cached bool like the sema/videoout flags.

* [HLE] Read guest C-strings without per-call buffer allocations

CpuContext.TryReadNullTerminatedUtf8 allocated a byte[capacity] and
issued one TryRead per byte for every string-argument import. It now
reads through a stack buffer (pooled above 512 bytes) in 128-byte bulk
chunks, falling back to per-byte reads only when a chunk touches an
unreadable range so a terminator sitting just before unmapped memory
still resolves exactly as before. The chunk bound also keeps the
overread past the terminator smaller than the old loop's worst case is
wide, so no fault can appear where the byte loop succeeded.

TryReadAsciiZ (dlsym/symbol resolution) drops its List<byte> + ToArray
round-trip for the same stack/pooled buffer, keeping the byte-by-byte
TryReadByteCompat reads because their Marshal.ReadByte fallback must
probe exactly up to the terminator. Only the final string is allocated
on either path now.

* [HLE] Replace blocking-wait closures with waiter continuation objects

Every wait that actually parked a guest thread allocated two capturing
lambdas (plus their display classes) for the scheduler's resume/wake
callbacks. RequestCurrentThreadBlock and the backend's blocked-thread
state now carry a single IGuestThreadBlockWaiter instead of the
Func<int>/Func<bool> pair: TryWake keeps the run-under-the-scheduler-
gate contract and Resume still produces the guest's RAX on the woken
thread. The waiter stays attached through the wake transition (the old
code nulled only the wake handler there) and is consumed at resume.

The existing waiter objects absorb the captured state as fields, so a
blocking wait now allocates exactly one object: SemaphoreWaiter,
PthreadMutexWaiter, and EventFlagWaiter implement the interface
directly, and the equeue, cond, and rwlock waits get small waiter
classes replacing their closures. Handler bodies delegate to the same
static methods with the same arguments as before; the untimed event
flag wait's mutable captured result becomes a field on its waiter.

* [HLE] Back pending event queues with a ring deque instead of LinkedList

LinkedList<KernelQueuedEvent> allocated a node object on every
non-coalesced enqueue — one per vblank/flip edge per registered queue,
60+ times a second in steady state. KernelEventDeque is a grow-only
ring buffer over a KernelQueuedEvent[] with the three operations the
queue actually uses (AddLast, RemoveFirst, find-and-update-in-place by
ident/filter), so steady-state enqueue/dequeue allocates nothing and
the coalescing update writes the struct back through an indexer instead
of a node reference. All accesses stay under _eventQueueGate, matching
the LinkedList usage it replaces.

* [HLE] Cap memcpy chunk iterations at the requested size, not the rented length

Address Copilot review: ArrayPool.Rent may return a larger array than
requested, so sizing each iteration by chunk.Length let the copy
granularity depend on pool bucketing internals instead of the intended
256 KB chunking. Behavior was already correct for any chunk size (each
iteration re-reads the source, and the overlap ordering is size-
independent), but the loop now mins against the requested chunkLength,
matching what memset already does.

* [HLE] Skip the flip/vblank snapshot rental when no events are registered

Address Copilot review: SignalVblank and SubmitFlip rented (and
returned) a pooled snapshot even with zero registrations — steady
per-frame pool traffic for games that never register flip events and
only poll flip status. Zero-count signals now skip the rental, the
copy, and the trigger loop entirely, which also retires the
Math.Max(count, 1) minimum-rent guard.
2026-07-15 12:57:40 +03:00
SamuelEzequias 6dacd59a08 [GUI] Add Portuguese (Portugal) translation (#197)
* [GUI] Add Portuguese (Portugal) translation

* [GUI] Add Portuguese (Portugal) translation
2026-07-15 12:56:25 +03:00
Spooks 9d88542efd Fix virtual memory allocation and access (#193)
* Fix virtual memory allocation and access

* Update test dependency lock file
2026-07-14 21:50:54 -06:00
StealUrKill 373100a6b0 Add 21 missing SysAbi exports and GUI Environment tab for SHARPEMU_* toggles (#189)
Fills NID gaps hit by PS5 titles during boot, controller setup, and
rendering, and surfaces the common runtime switches in the GUI. All
exports are additive (no behavior change to existing exports) and free of
NID and export-name collisions with upstream.

New export libraries:
- libSceBluetoothHid: Init/RegisterDevice/RegisterCallback success stubs so
  titles proceed past Bluetooth controller setup (opt-out via
  SHARPEMU_BTHID_UNAVAILABLE=1).
- libSceNpCppWebApi: Common::initialize no-op success; UE5 online titles
  abort PS5-component startup on a negative SCE error.

Additions to existing libraries:
- libScePad: scePadOpenExt (shared PadOpenCore, accepts special ports 1/2 and
  the ScePadOpenExtParam pointer), scePadClose, scePadGetExtControllerInformation.
- libSceVideoOut: sceVideoOutConfigureOutput, sceVideoOutInitializeOutputOptions.
- libSceAgc: DCB builders sceAgcDcbSetIndexCount, sceAgcDcbJump, DcbSetPredication,
  SetPacketPredication (emit valid skippable packets; full draw processing TODO).
- libSceAmpr: measure and write KernelEventQueueOnCompletion pair.
- libKernel: scePthreadGet/Setschedparam, sceKernelChmod (validate and accept;
  POSIX permission bits have no host equivalent on Windows).
- libSceNetCtl: sceNetCtlRegisterCallbackV6 (delegates to the v4 callback).
- libSceMouse: sceMouseInit.
- libSceUserService: sceUserServiceGetAgeLevel (adult, skips parental gates).

GUI: new Options Environment tab exposing common SHARPEMU_* switches as
toggles (BTHID_UNAVAILABLE, DISABLE_IMPORT_LOOP_GUARD, VK_VALIDATION,
DUMP_SPIRV, LOG_DIRECT_MEMORY, LOG_NP). Persisted in gui-settings.json and
applied to the emulator process environment at launch; localized with
English fallback.
2026-07-15 03:36:15 +03:00
Gutemberg Ribeiro f23161be9a Host platform abstraction layer for the execution engine (#181)
* [Host] Introduce host platform abstraction with IHostMemory

Add SharpEmu.HLE/Host with IHostPlatform/IHostMemory interfaces, neutral
page-protection/region enums, and a HostPlatform.Current factory that
resolves the Windows backend (or throws PlatformNotSupportedException on
other OSes, matching today's de-facto behavior). WindowsHostMemory wraps
the exact VirtualAlloc/VirtualFree/VirtualProtect/VirtualQuery calls used
across the engine today, with identical MEM_*/PAGE_* constants.

Migrate StubManager as the first consumer: its private kernel32 P/Invokes
and enums are replaced by IHostMemory calls that issue the same two
native operations (RWX commit+reserve of the PLT arena, release on
Dispose). No behavior change.

This is the first step toward supporting non-Windows hosts; subsequent
commits move the remaining direct P/Invokes in Core and Libs behind the
same seam.

* [Host] Route PhysicalVirtualMemory through IHostMemory

Replace the class's private VirtualAlloc/VirtualFree/VirtualProtect/
VirtualQuery P/Invokes with IHostMemory calls. Every site maps 1:1 onto
the exact native call it issued before: MEM_COMMIT|MEM_RESERVE ->
Allocate, MEM_RESERVE -> Reserve, fault-path commits -> Commit, and
MEM_RELEASE -> Free, with identical protection values produced by the
Windows backend.

IHostMemory gains ProtectRaw so the save/restore protection sequences in
TryWriteExclusive and TryTemporarilyProtectForRead round-trip the raw OS
protection word (including modifier bits the neutral enum cannot
represent) exactly as before. Raw PAGE_* constants remain only for the
internal region-classification helpers, which only ever see values this
class itself assigned.

The exact-address free-on-mismatch, lazy reserve-only threshold, prime
loop, and all trace strings are unchanged.

* [Host] Add IGuestAddressSpace and retire the reflection-based allocator lookup

Introduce IGuestAddressSpace in SharpEmu.HLE (fixed-address AllocateAt /
TryAllocateAtOrAbove and guest mprotect via TryProtect) with signatures
copied from PhysicalVirtualMemory, which now implements it. TryProtect
reproduces the read/write/execute decomposition that
KernelMemoryCompatExports.ResolveHostProtection performs, yielding the
same PAGE_* values through the Windows backend.

KernelVirtualRangeAllocator previously located AllocateAt via cached
MethodInfo reflection (because SharpEmu.Libs cannot see Core types) and
walked wrapper memories through an untyped 'Inner' property. Both are
now typed: ICpuMemoryWrapper exposes the decorated memory (implemented
by TrackedCpuMemory, whose Inner property already existed) and the
allocator type-tests for IGuestAddressSpace with the same bounded
unwrap depth. Failure paths keep the exact [LOADER][TRACE] strings.

* [Host] Move Kernel HLE memory exports off direct kernel32 P/Invokes

KernelMemoryCompatExports loses its private VirtualQuery/VirtualProtect/
VirtualAlloc/VirtualFree declarations and MemoryBasicInformation struct:

- Guest mprotect (sceKernelMprotect/sceKernelMtypeprotect) now routes
  through IGuestAddressSpace.TryProtect resolved from ctx.Memory. The
  orbis read/write/execute decomposition moves into a GuestPageProtection
  conversion whose mapping is value-identical to the removed
  ResolveHostProtection.
- The guarded libc heap and host-page accessibility checks go through
  IHostMemory (same commit+reserve/protect/free sequence; guard-page and
  protection-mask checks compare HostRegionInfo.RawProtection against the
  same PAGE_* literals as before).
- HostMemory is exposed as a property so merely loading the type never
  resolves the platform backend on non-Windows hosts.

KernelRuntimeCompatExports' RDTSC stub allocates its 16-byte RWX page via
IHostMemory.Allocate; the OperatingSystem.IsWindows() gate returning null
is unchanged.

* [Host] Abstract thread, TLS, and symbol primitives in the execution backend

Add IHostThreading (native TLS slots, current-thread id, affinity, raw
thread create/join, diagnostic register capture) and IHostSymbolResolver
(enum-keyed host function addresses baked into emitted stubs), with
Windows implementations wrapping the exact kernel32 calls the backend
made directly before.

DirectExecutionBackend takes an optional IHostPlatform (defaulting to
HostPlatform.Current) and routes every TlsAlloc/TlsFree/TlsSet/GetValue,
GetCurrentThreadId, SetThreadAffinityMask, GetModuleHandle/GetProcAddress
and the suspend+GetThreadContext diagnostic snapshot through it. The
snapshot moves wholesale into WindowsHostThreading (including the Win64
CONTEXT size/flags/offsets, which are Windows-specific by nature) and
returns a neutral HostCapturedRegisters.

NativeGuestExecutor resolves WaitForSingleObject/SetEvent/ExitThread via
the symbol resolver — the same addresses end up in the emitted run loop,
so stub bytes are unchanged — and creates/joins its raw worker thread
through IHostThreading with the same stack-reservation semantics. The
run-loop emitter itself does not move.

Marshal.GetLastWin32Error() in the affinity-failure log still observes
SetThreadAffinityMask's error because the wrapper makes no intervening
SetLastError call.

* [Host] Move fault handling and remaining backend memory ops behind the seam

Add IHostFaultHandling (handler-thunk creation, first-chance handler
install/remove, unhandled-filter set) with WindowsFaultHandling in a new
Cpu/Native/Windows/ folder. The exception-handler trampoline emitter
moves there whole — same pre-filtered NTSTATUS codes, same TEB gs:[8]/
gs:[0x10] stack-limit reads, same host-RSP TLS switch — parameterized
only by (managed callback, TLS slot, TlsGetValue address), which is
exactly what SetupExceptionHandler passed it before. Handler
installation order, the AddVectoredExceptionHandler(first=1) flag, the
SHARPEMU_DISABLE_RAW_HANDLER gate, and all install/teardown log strings
are unchanged.

Every remaining VirtualAlloc/VirtualProtect/VirtualFree/VirtualQuery/
FlushInstructionCache in the backend partials routes through IHostMemory
with 1:1 call mapping (RWX emit -> RX downgrade -> flush for stub
emission, reserve/commit for the PRT aperture and lazy-commit fault
path, raw-protection round-trips via ProtectRaw). HostRegionInfo gains
RawState/RawAllocationProtection so the lazy-commit trace lines and
protection-mask checks keep printing and comparing the exact native
values.

Windows semantics leaked as bare literals become named constants with
identical values: NTSTATUS codes (WindowsFaultCodes) and Win64 CONTEXT
byte offsets (Win64ContextOffsets, with the existing CTX_* constants
aliased to it and handler-local numeric offsets replaced by the names).

* [Host] Resolve the host platform explicitly at the composition root

SharpEmuRuntime.CreateDefault() now resolves HostPlatform.Current once
and passes it explicitly to PhysicalVirtualMemory and (via a new
optional CpuDispatcher parameter) to DirectExecutionBackend, replacing
the implicit default-argument fallbacks. On unsupported OSes boot now
fails at the root with PlatformNotSupportedException and a clear
message instead of on the first native call. A future Linux/macOS
backend plugs in by returning a different IHostPlatform here.

* [Host] Convert the platform backends to source-generated P/Invokes

Replace [DllImport] with [LibraryImport] in the four Windows backend
files added by this branch (WindowsHostMemory, WindowsHostThreading,
WindowsHostSymbolResolver, WindowsFaultHandling). Marshalling stubs are
now generated at compile time instead of JIT-emitted at runtime, which
fits the pre-JIT-everything boot model and keeps the backends
NativeAOT/trimming ready.

Interop stays zero-copy: all signatures are blittable, GetModuleHandleW
now pins the managed string via Utf16 marshalling instead of copying,
and GetProcAddress names marshal through a stack-allocated Utf8 buffer.
Implicit contracts become explicit where LibraryImport requires it:
TlsFree/TlsSetValue gain [MarshalAs(UnmanagedType.Bool)] (the 4-byte
Win32 BOOL DllImport assumed silently), and GetModuleHandle targets the
W entry point directly since LibraryImport never probes suffixes.

The CONTEXT snapshot buffer stays a NativeMemory allocation rather than
stackalloc: CONTEXT requires 16-byte alignment, now documented at the
call site. Native call sequences are unchanged.

* [Host] Address Copilot review: harden failure paths, honor injected platform

- Free the handler thunk page when the RX protection downgrade fails
  (the leak predates this branch, but the failure path is boot-fatal so
  releasing the page is unobservable).
- TraceThreadMode and the static diagnostics helpers now resolve host
  primitives through the backend bound to the current thread, falling
  back to HostPlatform.Current only when no run is active (identical on
  supported configs, honors injection everywhere a backend exists).
- HostPlatform.Create additionally requires an x64 process so native
  Windows ARM64 fails with the promised PlatformNotSupportedException
  instead of emitting x86-64 stubs into an ARM64 process.
2026-07-15 03:15:36 +03:00
Dafenx 081760be3f [AGC/Vulkan] Support multiple render targets (#149)
* [AGC] Support multiple typed pixel outputs

Emit dense float, uint, and sint fragment outputs for sparse guest MRT slots. Preserve disabled components across partial exports, validate dense host locations, and retain the single-output compiler overload for compatibility.

* [Vulkan] Execute translated draws with multiple color attachments

Carry every active color target and its effective shader/register write mask through one Vulkan draw. Add per-attachment blending, independentBlend negotiation, device/format validation, multi-attachment synchronization, and safe image recreation after in-flight work completes.

* [ShaderDump] Add MRT edge-case coverage

Cover sparse mixed-type outputs, partial exports, merged partial exports, independent blend layouts, eight attachments, and invalid host locations. Run the synthetic shader suite in CI.

---------

Co-authored-by: Dafenx <196083014+Dafenxz0@users.noreply.github.com>
2026-07-15 02:41:39 +03:00
José Luis Caravaca Carretero e604fb606d Fix pak size-collision that crashed Quake right after the intro demo (#187)
* [Tests] Add SharpEmu.Libs.Tests project

Introduce an xunit project for the HLE libs with a minimal ICpuMemory fake,
so library-level exports and helpers can be exercised without a live guest.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* [Ampr] Disambiguate pak size-collisions by read locality

PakDirectoryTracker resolves a sequential AMPR read (offset -1) back to an
absolute pak offset by matching the requested byte count against the PACK
directory. When several files share that byte count it took the first
unconsumed match in directory order, which mis-resolves out-of-order reads:
progs/h_ogre.mdl and bots/navigation/death32c.nav are both 0x3A34 bytes, and
death32c.nav sits earlier in the directory and is never read during Quake's
intro demo, so requesting h_ogre.mdl returned the nav file's bytes. The engine
then parsed "NAV2" as a brush model, failed the version check and aborted.

Pick the unconsumed same-size entry nearest the running read cursor instead.
id archives cluster related assets and the guest streams them with locality,
so this lands on the intended file; contiguous same-size runs (the
gfx/weapons/ww_*.lmp icons) still resolve in packed order.

Verified against a Quake dump: the abort is gone, h_ogre.mdl reads correctly,
and the intro demo reaches its main loop and renders instead of dying at the
error dialog.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 01:49:41 +03:00
José Luis Caravaca Carretero df53ff59d9 [Json] Implement sce::Json::Value and String (construct / set / destroy) (#169)
* [Json] Implement sce::Json::Value and Json::String construct/set/destroy

libSceJson previously only had the Initializer/MemAllocator setup path.
The Value and String classes themselves were entirely absent, so a
Prospero title that builds a JSON tree (Quake PPSA01880 does, to shape
a web-API request) hit unresolved imports and faulted on the call. The
imports it left unresolved right before its access violation are exactly
these Value ctors/setters and String ctor/dtor.

Model the Value/String payload host-side (JsonObjectHeap), keyed by the
guest `this` pointer, following the handle-shadow pattern already used
by Ngs2Exports. The guest object bytes are deliberately not written:
these objects are usually stack-allocated with an unknown real layout,
and writing a guessed layout risks smashing an adjacent stack canary
(the same hazard the AudioOut2 context-param note in this tree records).
Constructors and setters follow the Itanium ABI and return `this` in rax,
which is correct whether the real setter returns void or Value&.

Covered NIDs (complete-object C1/D1 variants, matching the observed
imports): Value(default/bool/long/ulong/double/ValueType/char*/String),
Value::~Value, Value::set(bool/long/ulong/double/ValueType/char*/String),
Value::clear, String(char*/default/copy), String::~String.

Only the payload the guest can reach through library methods is modelled;
direct guest reads of the object bytes are out of scope and would need
observed layout evidence.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* [Tests] Add SharpEmu.Libs.Tests covering the Json Value/String exports

First test project for SharpEmu.Libs (xunit), the SharpEmu.Libs.Tests
layout the maintainer already agreed to in issue #36.

- A FakeCpuMemory (single contiguous region) drives the exports at the
  CpuContext level with no live guest.
- Direct-call tests: ctor/setter round-trips for bool/int/uint/double
  (read from xmm0)/char*/String/ValueType, destructor cleanup, and the
  graceful-degradation paths (missing String shadow and a faulting char*
  pointer both fall back to the empty string instead of throwing).
- Registration test: a real ModuleManager scans SharpEmu.Libs and the
  nine NIDs Quake left unresolved now resolve to the libSceJson exports
  and dispatch cleanly (returns `this` in rax).

InternalsVisibleTo exposes JsonObjectHeap to the test assembly. The test
project's packages.lock.json is committed for CI locked-mode restore;
CI does not run tests yet, left as a maintainer decision.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* [Json] Add Initializer::setGlobalNullAccessCallback

Quake calls it during kexPSNWebAPI::Initialize and treats the
not-found error as fatal for the whole Np Web API bring-up. Store the
guest hook (never invoked by this HLE: shadows degrade to defaults
instead of dereferencing missing members) and return success.

Verified against the dump: the "setGlobalNullAccessCallback failed
(0x80020002)" line is gone and kexPSNWebAPI::Initialize now logs
"Np Web API Initialized"; the next blockers are sceNpAuthCreateRequest
and sceUserServiceInitialize ordering, outside libSceJson.

Also pins both Json test classes to one xunit collection: they share
JsonObjectHeap statics and parallel class execution raced ResetForTests
against a running test.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 01:49:33 +03:00
Randomuser8219 4c35831cb8 Add code -1073741819 as an emulation error (#188)
There's currently games that crash on this code due to emulation errors, so it'd make sense to add this error code as an emulation error.
2026-07-15 01:47:44 +03:00
miles 90fdd20f9a Create nl.json (#186) 2026-07-15 01:36:38 +03:00
Mike Saito ae5ef0abe7 Add SaveData transaction and NP UDS layout HLE stubs (#168)
* Add SaveData transaction and NP UDS layout HLE stubs

Wire Prepare, Commit, and Umount2 for implicit save transactions,
unregister guest mounts on Umount2, and add NP UDS CreateEvent,
DestroyEvent, and EventPropertyObjectSetString for layout-load imports.

* Add NP UDS SetArray and PostEvent layout HLE stubs

Add sceNpUniversalDataSystemEventPropertyObjectSetArray and
sceNpUniversalDataSystemPostEvent for layout-load imports on PPSA02929.
2026-07-15 01:34:59 +03:00
Mike Saito 5e2c21edf1 Fix historic SysAbi exports bound to wrong symbol names (#167)
Move KMcEa+rHsIo from libKernel MapMemory mislabel to sceAvPlayerAddSource.
Align WV1GwM32NgY ExportName with sceNpWebApi2PushEventCreateHandle. Behavior unchanged.
2026-07-15 01:34:27 +03:00
Deeptanshu Lal 3fb9d4db1c [Tools] Fix ShaderDump reflection invoke against new optional parameters (#166)
TryCompileVertexShader gained an optional scalarRegisterBufferIndex
parameter (#156), and reflection Invoke does not apply C# default
parameter values, so ShaderDump crashed with
TargetParameterCountException. Pad trailing optional parameters with
Type.Missing under BindingFlags.OptionalParamBinding so the declared
defaults are used; only a new required parameter now needs a tool
update, and that fails with a named error instead of a crash.

Verified: all five programs behave as expected (exit 0), all eight
emitted blobs pass spirv-val --target-env vulkan1.3.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 01:34:08 +03:00
José Luis Caravaca Carretero de13735972 [CommonDialog] Fix dialog state machine and add MsgDialog progress-bar exports (#163)
Rework the sceMsgDialog and sceSaveDataDialog HLE state machines so the full
Initialize -> Open -> poll -> GetResult -> Close/Terminate lifecycle honors the
common-dialog contract, and add the three missing sceMsgDialogProgressBar* exports.

- Fix an unreachable close path: sceSaveDataDialogClose already did a
  RUNNING -> FINISHED compare-exchange, but Open jumped straight to FINISHED, so
  RUNNING never existed and Close could only return NOT_RUNNING. Open now enters
  RUNNING and the first status poll advances it to FINISHED. Same model applied to
  sceMsgDialog.
- Return the real SCE_COMMON_DIALOG_ERROR_* codes (0x80B8xxxx) from sceMsgDialog*
  instead of emulator-internal result codes, with the missing argument/state guards
  (ARG_NULL, NOT_INITIALIZED, BUSY, NOT_FINISHED, NOT_RUNNING).
- GetResult reports buttonId = 1 (affirmative) instead of 0, the invalid sentinel a
  yes/no prompt could mis-branch on.
- Add sceMsgDialogProgressBarSetValue, sceMsgDialogProgressBarInc and
  sceMsgDialogProgressBarSetMsg (NIDs wTpfglkmv34, Gc5k1qcK4fs, 6H-71OdrpXM), gated
  on the service being initialized.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 01:33:48 +03:00
AlexC fc0efca297 Fixed deutch language file, it had invalid syntax (#162) 2026-07-15 01:31:54 +03:00
anesr5 2a9a261913 loader: support ps5 SELF and validate ELF signatures (#157)
Co-authored-by: anes <anesrachedi@outlook.fr>
2026-07-15 01:31:16 +03:00
Mike Saito 290f5fd3d7 Add SysAbi ExportName name2nid check script (#152)
* Add SysAbi ExportName name2nid check script

* Make SysAbi ExportName check green on tip with catalog skips and one Np rename
2026-07-15 01:29:23 +03:00
tensorcrush c06c70cad7 [Aerolib] Add ulobjmgr and NpEAAccess symbol names (#150)
Resolves _sceUlobjmgrRegisterObject (BG26hBGiNlw) and
_sceUlobjmgrUnregisterObject (Smf+fUNblPc), reported as unresolved by
testers, plus four sceNpEAAccess exports. Names taken from shadPS4's
NID tables and each verified by recomputing the NID with the repo's
name2nid derivation before inclusion. aerolib.bin regenerated with
scripts/generate_aerolib_binary.py.

Co-authored-by: tensorcrush <tensorcrush@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 01:28:17 +03:00
Deeptanshu Lal 5e54250752 [Tools] Add GPU conformance executor for dumped shader blobs (#127)
SharpEmu.Tools.GpuConformance executes the exec-cs.spv blob produced by
SharpEmu.Tools.ShaderDump on a real Vulkan device (preferring a discrete
GPU) and compares every word of the 64-byte storage buffer against
CPU-computed expectations, bit for bit. Creating the compute pipeline
doubles as a driver-acceptance check for SharpEmu's emitted SPIR-V.

The checks cover the three ALU results, the store attempted with EXEC=0
(its destination must keep the sentinel), the store after EXEC is
restored, and all trailing sentinel words. Any mismatch counts toward the
failure total and makes the tool exit non-zero.

Verified on an RTX 3060 Laptop GPU (NVIDIA) with all values matching, and
the failure path verified to exit 1 by running a non-storing blob.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-15 01:26:18 +03:00
AlexC be6a6a5935 [GUI] New box "About" in configuration (#165)
* Added about tab with github and discord

* Added discord & github svgs and svg support

* Changed svg to pngs and localization text in english & spanish
2026-07-15 00:59:13 +03:00
Mike Saito caf859cc52 Fix guest shutdown when VideoOut window is closed (#184)
Propagate Silk window close to runtime teardown so audio and CPU workers stop instead of continuing after the presentation window is dismissed.
2026-07-15 00:52:01 +03:00
brbrhuehue-matrix d2f3511002 Add Brazilian Portuguese translation (#153) 2026-07-15 00:42:44 +03:00
Nolan 90a5d5176f Add Korean (ko-KR) localization (#154) 2026-07-15 00:42:31 +03:00
Nolan 28a43e09c7 Add Japanese (ja) localization (#160) 2026-07-15 00:42:17 +03:00
AlexC 093cfa1f3e Fallback to english if it doesnt find the string in current language (#161) 2026-07-15 00:42:10 +03:00
Spooks d8397b022e Performance Improvements and Optimization Tweaks (#156)
* Improve Gen5 rendering performance and compatibility

* Pin .NET SDK for locked restore

---------

Co-authored-by: Spooks4576 <Spooks4576@users.noreply.github.com>
2026-07-14 20:22:52 +03:00
anesr5 85cc2b9892 added french support (#147)
Co-authored-by: anes <anesrachedi@outlook.fr>
2026-07-14 18:06:57 +03:00
AlexC 293194c40b [GUI] Add Spanish language (#148)
* Added localization to spanish language

* Changed Options.Strict.Desc because i didnt like the way i localized it first
2026-07-14 18:06:48 +03:00
tensorcrush 1f09de8896 [AGC] Complete gfx10 v_cmpx_f32 decode and emit ordered/unordered float compares (#122)
* [AGC] Complete gfx10 v_cmpx_f32 decode and emit ordered/unordered float compares

Add the missing v_cmpx_*_f32 VOPC decode entries (0x17-0x1C, 0x1F) and
emission for the ordered/unordered predicates: nlg maps to OpFUnordEqual,
while o/u are lowered from OpIsNan (unordered = isnan(a) || isnan(b),
ordered = !unordered) because SPIR-V's OpOrdered/OpUnordered require the
Kernel capability and are invalid in Vulkan shader modules.

Opcode numbers cross-checked against LLVM's llvm-mc regression tests
(llvm/test/MC/AMDGPU/gfx10_asm_vopc.s, gfx10_asm_vopcx.s); emitted
lowering validated with spirv-val --target-env vulkan1.1.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* [AGC] Write VCC only for non-X vector compares

On gfx10 the VCmpx encodings have no sdst and define EXEC only, so the
unconditional VCC store clobbered VCC on every VCmpx. Move the VCC store
to the non-X path; EXEC keeps the existing old-EXEC & condition update.

Addresses review feedback on #122.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: tensorcrush <tensorcrush@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 18:05:33 +03:00
442 changed files with 101492 additions and 13600 deletions
Binary file not shown.

After

Width:  |  Height:  |  Size: 190 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 229 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 225 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 86 KiB

+27
View File
@@ -0,0 +1,27 @@
## Before submitting
Please read our contribution guidelines before opening a pull request:
➡️ [**CONTRIBUTING.md**](https://github.com/sharpemu/sharpemu/blob/main/CONTRIBUTING.md)
By opening this pull request, you confirm that you have read and agree to follow the contribution guidelines.
## Testing
If applicable, list the game(s) you tested and briefly describe the results.
Example:
- Demon's Souls (PPSA01341) Boots to splash screen.
- Dreaming Sarah Save/load works correctly.
If your changes do not affect runtime behavior (e.g. documentation, tooling, CI), write `N/A`.
## Checklist
By submitting this pull request, you confirm that:
- [ ] I have read and followed `CONTRIBUTING.md`.
- [ ] I tested my changes or marked the testing section as `N/A`.
- [ ] I wrote this pull request description myself and did not paste AI-generated explanations.
- [ ] I listed the game(s) I tested (or marked the testing section as `N/A`).
+31
View File
@@ -0,0 +1,31 @@
# Copyright (C) 2026 SharpEmu Emulator Project
# SPDX-License-Identifier: GPL-2.0-or-later
# Tells the website repo to rebuild when a release is published, so
# sharpemu.app/downloads lists the new build within a minute.
name: Notify website
on:
release:
types: [published, released, edited, deleted]
workflow_dispatch:
jobs:
dispatch:
runs-on: ubuntu-latest
steps:
- name: Trigger sharpemu-site rebuild
env:
TOKEN: ${{ secrets.SITE_DISPATCH_TOKEN }}
run: |
if [ -z "$TOKEN" ]; then
echo "SITE_DISPATCH_TOKEN is not set — skipping website rebuild."
exit 0
fi
curl -fsS -X POST \
-H "Authorization: Bearer $TOKEN" \
-H "Accept: application/vnd.github+json" \
-H "X-GitHub-Api-Version: 2022-11-28" \
https://api.github.com/repos/sharpemu/sharpemu-site/dispatches \
-d '{"event_type":"release-published"}'
echo "Website rebuild requested."
+120
View File
@@ -0,0 +1,120 @@
# Copyright (C) 2026 SharpEmu Emulator Project
# SPDX-License-Identifier: GPL-2.0-or-later
# Posts (and keeps updated) a single PR comment linking the per-platform build
# artifacts once "Build and Release" finishes.
#
# This is a workflow_run workflow on purpose: PRs from forks run "Build and
# Release" with a read-only GITHUB_TOKEN and cannot comment. workflow_run runs
# afterwards in the base-repo context with a read-write token and does not check
# out untrusted fork code, so it can comment safely. Because of that, GitHub only
# triggers it from the copy on the default branch — it does nothing until merged
# to main.
name: PR Build Links
on:
workflow_run:
workflows: ["Build and Release"]
types:
- completed
permissions:
contents: read
actions: read
pull-requests: write
jobs:
comment:
name: Post artifact links
runs-on: ubuntu-latest
# Only for successful PR builds — artifacts exist only when the build passed.
if: >-
github.event.workflow_run.event == 'pull_request' &&
github.event.workflow_run.conclusion == 'success'
steps:
- name: Post or update the artifact-links comment
uses: actions/github-script@v7
with:
script: |
const run = context.payload.workflow_run;
const { owner, repo } = context.repo;
// Resolve the PR. Same-repo PRs are in workflow_run.pull_requests, but
// fork PRs leave it empty AND their head commit is not on a base-repo
// branch, so listPullRequestsAssociatedWithCommit misses them too. Match
// the open PR by its head "owner:branch" instead, which works for forks.
let prNumber;
if (run.pull_requests && run.pull_requests.length > 0) {
prNumber = run.pull_requests[0].number;
} else {
let candidates = [];
const headOwner = run.head_repository && run.head_repository.owner
? run.head_repository.owner.login
: null;
if (headOwner && run.head_branch) {
const byHead = await github.rest.pulls.list({
owner, repo, state: 'open',
head: `${headOwner}:${run.head_branch}`, per_page: 10,
});
candidates = byHead.data;
}
if (candidates.length === 0) {
const byCommit = await github.rest.repos.listPullRequestsAssociatedWithCommit({
owner, repo, commit_sha: run.head_sha,
});
candidates = byCommit.data.filter(pr => pr.state === 'open');
}
if (candidates.length === 0) {
core.info('No open PR for this build; nothing to comment.');
return;
}
prNumber = candidates[0].number;
}
// Collect the per-platform artifacts the build produced.
const artifacts = await github.paginate(
github.rest.actions.listWorkflowRunArtifacts,
{ owner, repo, run_id: run.id, per_page: 100 },
);
const platforms = [
{ key: 'win-x64', label: 'Windows (win-x64)' },
{ key: 'linux-x64', label: 'Linux (linux-x64)' },
{ key: 'osx-x64', label: 'macOS (osx-x64)' },
];
const rows = [];
for (const platform of platforms) {
const artifact = artifacts.find(a => a.name.includes(platform.key));
if (!artifact) {
continue;
}
const url = `https://github.com/${owner}/${repo}/actions/runs/${run.id}/artifacts/${artifact.id}`;
rows.push(`| ${platform.label} | [\`${artifact.name}\`](${url}) |`);
}
if (rows.length === 0) {
core.info('No platform artifacts on this run; nothing to comment.');
return;
}
const marker = '<!-- pr-build-links -->';
const body = [
marker,
`### 📦 Build artifacts — \`${run.head_sha.substring(0, 7)}\``,
'',
'| Platform | Download |',
'| --- | --- |',
...rows,
'',
`From [build run #${run.run_number}](${run.html_url}). ` +
'Downloads require a GitHub login and expire after 90 days.',
].join('\n');
// Upsert one comment so repeated builds refresh it in place.
const comments = await github.paginate(github.rest.issues.listComments, {
owner, repo, issue_number: prNumber, per_page: 100,
});
const existing = comments.find(c => c.body && c.body.includes(marker));
if (existing) {
await github.rest.issues.updateComment({ owner, repo, comment_id: existing.id, body });
} else {
await github.rest.issues.createComment({ owner, repo, issue_number: prNumber, body });
}
+164 -38
View File
@@ -7,6 +7,8 @@ on:
push:
branches:
- "**"
tags:
- "v*"
paths-ignore:
- "**/*.md"
- "**/*.png"
@@ -34,29 +36,36 @@ jobs:
name: Init
runs-on: ubuntu-latest
outputs:
archive-name: ${{ steps.vars.outputs.archive-name }}
artifact-name: ${{ steps.vars.outputs.artifact-name }}
release-name: ${{ steps.vars.outputs.release-name }}
release-tag: ${{ steps.vars.outputs.release-tag }}
safe-ref: ${{ steps.vars.outputs.safe-ref }}
short-sha: ${{ steps.vars.outputs.short-sha }}
version: ${{ steps.vars.outputs.version }}
steps:
- name: Checkout repository
uses: actions/checkout@v6
- name: Compute workflow variables
id: vars
shell: bash
run: |
short_sha="${GITHUB_SHA::7}"
safe_ref="$(echo "${GITHUB_REF_NAME}" | tr '[:upper:]' '[:lower:]' | sed 's#[^a-z0-9._-]#-#g')"
archive_name="sharpemu-win64-${short_sha}.zip"
artifact_name="sharpemu-win64-${short_sha}"
release_tag="win64-${safe_ref}-${short_sha}"
release_name="SharpEmu win64 ${short_sha}"
version="$(python3 -c 'import sys, xml.etree.ElementTree as ET; version = ET.parse("Directory.Build.props").getroot().findtext(".//SharpEmuVersion"); sys.exit("Directory.Build.props is missing SharpEmuVersion") if not version or not version.strip() else print(version.strip())')"
release_tag="v${version}"
release_name="SharpEmu v${version}"
if [ "${GITHUB_REF_TYPE}" = "tag" ] && [ "${GITHUB_REF_NAME}" != "${release_tag}" ]; then
echo "Release tag ${GITHUB_REF_NAME} does not match project version ${release_tag}." >&2
exit 1
fi
{
echo "short-sha=${short_sha}"
echo "archive-name=${archive_name}"
echo "artifact-name=${artifact_name}"
echo "safe-ref=${safe_ref}"
echo "release-tag=${release_tag}"
echo "release-name=${release_name}"
echo "version=${version}"
} >> "$GITHUB_OUTPUT"
reuse:
@@ -80,7 +89,6 @@ jobs:
DOTNET_NOLOGO: true
NUGET_PACKAGES: ${{ github.workspace }}\.nuget\packages
PUBLISH_DIR: ${{ github.workspace }}\artifacts\publish\win-x64
RELEASE_DIR: ${{ github.workspace }}\artifacts\release
steps:
- name: Checkout repository
uses: actions/checkout@v6
@@ -92,69 +100,187 @@ jobs:
cache: true
cache-dependency-path: |
Directory.Packages.props
src/**/packages.lock.json
Directory.Build.props
- name: Restore solution
run: dotnet restore SharpEmu.slnx --locked-mode
run: dotnet restore SharpEmu.slnx
- name: Build solution
run: dotnet build SharpEmu.slnx -c Release --no-restore
# Runs every test project in the solution; a test failure fails the build, as
# does an aerolib.bin generation failure in the build step above (the MSBuild
# task logs an error and returns false).
- name: Run tests
run: dotnet test SharpEmu.slnx -c Release --no-build --verbosity normal
- name: Validate synthetic shaders
run: dotnet run --project tools/SharpEmu.Tools.ShaderDump/SharpEmu.Tools.ShaderDump.csproj -c Release -- artifacts/shader-dump
- name: Publish win-x64 CLI
run: dotnet publish src/SharpEmu.CLI/SharpEmu.CLI.csproj -c Release -r win-x64 --self-contained true --no-restore -p:PublishDir="${env:PUBLISH_DIR}"
- name: Create release archive
run: |
New-Item -ItemType Directory -Path $env:RELEASE_DIR -Force | Out-Null
$archivePath = Join-Path $env:RELEASE_DIR "${{ needs.init.outputs.archive-name }}"
if (Test-Path $archivePath) {
Remove-Item $archivePath -Force
}
Compress-Archive -Path (Join-Path $env:PUBLISH_DIR '*') -DestinationPath $archivePath -CompressionLevel Optimal
- name: Upload build artifact
uses: actions/upload-artifact@v7
with:
name: ${{ needs.init.outputs.artifact-name }}
path: ${{ env.RELEASE_DIR }}\${{ needs.init.outputs.archive-name }}
name: sharpemu-win-x64-${{ needs.init.outputs.short-sha }}
path: ${{ env.PUBLISH_DIR }}
if-no-files-found: error
include-hidden-files: true
build-posix:
name: Build ${{ matrix.rid }}
needs:
- init
- reuse
runs-on: ${{ matrix.os }}
strategy:
fail-fast: false
matrix:
include:
- os: ubuntu-latest
rid: linux-x64
- os: macos-latest
rid: osx-x64
env:
DOTNET_NOLOGO: true
NUGET_PACKAGES: ${{ github.workspace }}/.nuget/packages
PUBLISH_DIR: ${{ github.workspace }}/artifacts/publish/${{ matrix.rid }}
SPIRV_HEADERS_COMMIT: ad9184e76a66b1001c29db9b0a3e87f646c64de0
# SpirvModuleBuilder emits SPIR-V 1.5 and VulkanVideoPresenter requests Vulkan 1.2.
SPIRV_TARGET_ENV: vulkan1.2
SPIRV_TOOLS_COMMIT: 0539c81f69a3daeb706fd3477dca61435b475156
SPIRV_TOOLS_VERSION: v2026.2
steps:
- name: Checkout repository
uses: actions/checkout@v6
- name: Setup .NET SDK
uses: actions/setup-dotnet@v5
with:
dotnet-version: 10.0.103
cache: true
cache-dependency-path: |
Directory.Packages.props
Directory.Build.props
- name: Restore solution
run: dotnet restore SharpEmu.slnx
- name: Build solution
run: dotnet build SharpEmu.slnx -c Release --no-restore
# Also runs on Linux/macOS: the aerolib.bin MSBuild task does platform-sensitive
# path handling, so a cross-platform generation regression surfaces here.
- name: Run tests
run: dotnet test SharpEmu.slnx -c Release --no-build --verbosity normal
- name: Build pinned SPIRV-Tools
if: matrix.rid == 'linux-x64'
run: |
git clone --no-checkout --filter=blob:none https://github.com/KhronosGroup/SPIRV-Tools.git "$RUNNER_TEMP/spirv-tools"
git -C "$RUNNER_TEMP/spirv-tools" checkout --detach "$SPIRV_TOOLS_COMMIT"
test "$(git -C "$RUNNER_TEMP/spirv-tools" rev-parse HEAD)" = "$SPIRV_TOOLS_COMMIT"
git clone --no-checkout --filter=blob:none https://github.com/KhronosGroup/SPIRV-Headers.git "$RUNNER_TEMP/spirv-tools/external/spirv-headers"
git -C "$RUNNER_TEMP/spirv-tools/external/spirv-headers" checkout --detach "$SPIRV_HEADERS_COMMIT"
test "$(git -C "$RUNNER_TEMP/spirv-tools/external/spirv-headers" rev-parse HEAD)" = "$SPIRV_HEADERS_COMMIT"
cmake -S "$RUNNER_TEMP/spirv-tools" -B "$RUNNER_TEMP/spirv-tools-build" \
-G Ninja \
-DCMAKE_BUILD_TYPE=Release \
-DSPIRV_SKIP_TESTS=ON \
-DSPIRV_WERROR=OFF
cmake --build "$RUNNER_TEMP/spirv-tools-build" --target spirv-val
- name: Generate and validate synthetic SPIR-V
if: matrix.rid == 'linux-x64'
run: |
dotnet run --project tools/SharpEmu.Tools.ShaderDump/SharpEmu.Tools.ShaderDump.csproj -c Release -- artifacts/shader-dump
scripts/validate-synthetic-spirv.sh \
"$RUNNER_TEMP/spirv-tools-build/tools/spirv-val" \
"$SPIRV_TOOLS_VERSION" \
"$SPIRV_TARGET_ENV" \
artifacts/shader-dump
- name: Publish ${{ matrix.rid }} CLI
run: dotnet publish src/SharpEmu.CLI/SharpEmu.CLI.csproj -c Release -r ${{ matrix.rid }} --self-contained true --no-restore -p:PublishDir="$PUBLISH_DIR"
- name: Stage MoltenVK next to the build
if: matrix.rid == 'osx-x64'
run: scripts/fetch-macos-moltenvk.sh "$PUBLISH_DIR"
- name: Upload build artifact
uses: actions/upload-artifact@v7
with:
name: sharpemu-${{ matrix.rid }}-${{ needs.init.outputs.short-sha }}
path: ${{ env.PUBLISH_DIR }}
if-no-files-found: error
include-hidden-files: true
release:
name: Publish GitHub Release
needs:
- init
- build
if: github.event_name == 'workflow_dispatch' || (github.event_name == 'push' && github.ref == 'refs/heads/main')
- build-posix
# Versioned releases are immutable and tag-driven. Branch and manual runs
# still produce Actions artifacts without modifying a published release.
if: github.event_name == 'push' && startsWith(github.ref, 'refs/tags/v')
runs-on: ubuntu-latest
permissions:
contents: write
steps:
- name: Download build artifact
- name: Download build artifacts
uses: actions/download-artifact@v8
with:
name: ${{ needs.init.outputs.artifact-name }}
path: release
- name: Create or update release
- name: Package release assets
shell: bash
env:
SHORT_SHA: ${{ needs.init.outputs.short-sha }}
VERSION: ${{ needs.init.outputs.version }}
run: |
set -euo pipefail
win_dir="release/sharpemu-win-x64-${SHORT_SHA}"
linux_dir="release/sharpemu-linux-x64-${SHORT_SHA}"
macos_dir="release/sharpemu-osx-x64-${SHORT_SHA}"
for package_dir in "${win_dir}" "${linux_dir}" "${macos_dir}"; do
test -d "${package_dir}"
done
mkdir -p release-assets
(cd "${win_dir}" && zip -q -r "../../release-assets/sharpemu-${VERSION}-win-x64.zip" .)
chmod +x "${linux_dir}/SharpEmu" "${macos_dir}/SharpEmu"
tar -czf "release-assets/sharpemu-${VERSION}-linux-x64.tar.gz" -C "${linux_dir}" .
tar -czf "release-assets/sharpemu-${VERSION}-osx-x64.tar.gz" -C "${macos_dir}" .
- name: Create release
shell: bash
env:
ARCHIVE_NAME: ${{ needs.init.outputs.archive-name }}
GH_REPO: ${{ github.repository }}
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
RELEASE_NAME: ${{ needs.init.outputs.release-name }}
RELEASE_TAG: ${{ needs.init.outputs.release-tag }}
VERSION: ${{ needs.init.outputs.version }}
run: |
asset_path="release/${ARCHIVE_NAME}"
notes="Automated Windows build for commit ${GITHUB_SHA}."
mapfile -t assets < <(find release-assets -maxdepth 1 -type f \( -name '*.zip' -o -name '*.tar.gz' \) | sort)
if [ "${#assets[@]}" -ne 3 ]; then
echo "Expected 3 release assets, found ${#assets[@]}." >&2
exit 1
fi
notes="Automated SharpEmu v${VERSION} build for commit ${GITHUB_SHA}."
if gh release view "${RELEASE_TAG}" >/dev/null 2>&1; then
gh release upload "${RELEASE_TAG}" "${asset_path}" --clobber
gh release edit "${RELEASE_TAG}" --title "${RELEASE_NAME}" --notes "${notes}"
else
gh release create "${RELEASE_TAG}" "${asset_path}" \
--title "${RELEASE_NAME}" \
--notes "${notes}" \
--target "${GITHUB_SHA}"
echo "Release ${RELEASE_TAG} already exists and will not be modified." >&2
exit 1
fi
gh release create "${RELEASE_TAG}" "${assets[@]}" \
--verify-tag \
--title "${RELEASE_NAME}" \
--notes "${notes}"
+3
View File
@@ -32,6 +32,8 @@ packages/
.nuget/
.dotnet-home/
.cache/
__pycache__/
*.py[cod]
.DS_Store
Thumbs.db
@@ -40,3 +42,4 @@ ehthumbs.db
.vs/
.idea/
.vscode/
+24
View File
@@ -5,6 +5,11 @@ SPDX-License-Identifier: GPL-2.0-or-later
# Contributing
> [!IMPORTANT]
> The pull request template is mandatory.
>
> Pull requests that do not follow the template or leave the required checklist incomplete will be closed without review, even if the proposed code is technically correct or beneficial. Please review these contribution guidelines before submitting a pull request.
Contributions are always welcome!
Before opening a pull request, please keep the following in mind:
@@ -21,6 +26,25 @@ Before opening a pull request, please keep the following in mind:
If you're unsure about a design decision, feel free to open a discussion or draft PR first.
## Pull Request Expectations
Pull requests should provide real, observable emulator behavior rather than only suppressing errors or unresolved imports.
Changes that only return success, zero, or fabricated handles without implementing the expected state, output, or side effects will generally not be accepted. Functions that create resources, write output structures, register callbacks, or expose runtime state should model the behavior required by the guest.
When applicable, PRs should include:
- The affected game or application.
- Relevant logs or failing imports.
- Behavior before and after the change.
- Real game testing and known limitations.
Avoid submitting large collections of speculative NIDs or unrelated exports. Keep each PR focused on one problem or a closely related set of changes.
Large architectural changes should be discussed with the maintainers before implementation. Contributors are encouraged to ask first when they are uncertain whether a proposed direction fits the project.
Opening a PR does not guarantee that it will be merged. Maintainers evaluate changes based on correctness, evidence, testing, scope, maintenance cost, and the long-term direction of the project.
## AI-Assisted Contributions
AI-assisted development is welcome and may be used for research, reverse engineering, code generation, or documentation.
+9 -1
View File
@@ -9,10 +9,18 @@ SPDX-License-Identifier: GPL-2.0-or-later
<ImplicitUsings>enable</ImplicitUsings>
<Nullable>enable</Nullable>
<GenerateDocumentationFile>true</GenerateDocumentationFile>
<RestorePackagesWithLockFile>true</RestorePackagesWithLockFile>
<SharpEmuVersion>0.0.2-beta.4</SharpEmuVersion>
<Version>$(SharpEmuVersion)</Version>
<RepoRoot>$([MSBuild]::NormalizeDirectory('$(MSBuildThisFileDirectory)'))</RepoRoot>
<_HostRidOSPrefix Condition="'$(RuntimeIdentifier)' == '' And '$(MSBuildProjectName)' == 'SharpEmu.CLI' And $([MSBuild]::IsOSPlatform('Windows'))">win</_HostRidOSPrefix>
<_HostRidOSPrefix Condition="'$(RuntimeIdentifier)' == '' And '$(MSBuildProjectName)' == 'SharpEmu.CLI' And '$(_HostRidOSPrefix)' == '' And $([MSBuild]::IsOSPlatform('Linux'))">linux</_HostRidOSPrefix>
<_HostRidOSPrefix Condition="'$(RuntimeIdentifier)' == '' And '$(MSBuildProjectName)' == 'SharpEmu.CLI' And '$(_HostRidOSPrefix)' == '' And $([MSBuild]::IsOSPlatform('OSX'))">osx</_HostRidOSPrefix>
<_HostRidArch Condition="'$(_HostRidOSPrefix)' != '' And '$([System.Runtime.InteropServices.RuntimeInformation]::ProcessArchitecture)' == 'Arm64'">arm64</_HostRidArch>
<_HostRidArch Condition="'$(_HostRidOSPrefix)' != '' And '$(_HostRidArch)' == ''">x64</_HostRidArch>
<RuntimeIdentifier Condition="'$(_HostRidOSPrefix)' != ''">$(_HostRidOSPrefix)-$(_HostRidArch)</RuntimeIdentifier>
<BaseIntermediateOutputPath>$(RepoRoot)artifacts/obj/$(MSBuildProjectName)/</BaseIntermediateOutputPath>
<BaseOutputPath>$(RepoRoot)artifacts/bin/</BaseOutputPath>
+9 -1
View File
@@ -11,12 +11,20 @@ SPDX-License-Identifier: GPL-2.0-or-later
<PackageVersion Include="Avalonia.Desktop" Version="11.3.18" />
<PackageVersion Include="Avalonia.Fonts.Inter" Version="11.3.18" />
<PackageVersion Include="Avalonia.Themes.Fluent" Version="11.3.18" />
<PackageVersion Include="FFmpeg.AutoGen" Version="7.1.1" />
<PackageVersion Include="Iced" Version="1.21.0" />
<PackageVersion Include="Microsoft.Build.Framework" Version="17.14.8" />
<PackageVersion Include="Microsoft.CodeAnalysis.Analyzers" Version="3.11.0" />
<PackageVersion Include="Microsoft.CodeAnalysis.CSharp" Version="4.12.0" />
<PackageVersion Include="Microsoft.NET.Test.Sdk" Version="17.14.1" />
<PackageVersion Include="Silk.NET.Input" Version="2.23.0" />
<PackageVersion Include="Silk.NET.Vulkan" Version="2.23.0" />
<PackageVersion Include="Silk.NET.Vulkan.Extensions.EXT" Version="2.23.0" />
<PackageVersion Include="Silk.NET.Vulkan.Extensions.KHR" Version="2.23.0" />
<PackageVersion Include="Silk.NET.Windowing" Version="2.23.0" />
<!-- Transitive of Avalonia.Desktop; pinned to fix GHSA-xrw6-gwf8-vvr9 -->
<PackageVersion Include="Tmds.DBus.Protocol" Version="0.21.3" />
<PackageVersion Include="xunit" Version="2.9.3" />
<PackageVersion Include="xunit.runner.visualstudio" Version="3.1.1" />
</ItemGroup>
</Project>
</Project>
+36 -23
View File
@@ -25,8 +25,11 @@ SPDX-License-Identifier: GPL-2.0-or-later
---
> [!WARNING]
> Currently the primary development target is Windows.
> [!NOTE]
> SharpEmu supports Windows x64, Linux x64, and macOS x64. Apple Silicon Macs
> can run the macOS x64 build through Rosetta 2, and Windows on ARM devices
> (e.g. Snapdragon) can run the Windows x64 build through Windows' built-in
> x64 emulation.
> [!WARNING]
> SharpEmu is an experimental PS5 emulator developed from scratch in C#. The current focus is on accuracy and infrastructure setup rather than game-specific compatibility.
@@ -40,6 +43,16 @@ This project is developed purely for research and educational purposes. There ar
SharpEmu focuses exclusively on the PlayStation 5.
Our goal is **not** to emulate PS4 games, as there is already an excellent emulator dedicated to that platform: **ShadPS4**.
## Games Tested
| Demons Souls Remake | Dreaming Sarah |
| :-----------------------------------------------------------: | :--------------------------------------------------------------------------------------------: |
| ![Bloodborne screenshot](./.github/images/demons-souls.jpg) | ![Dreaming Sarah](./.github/images/dreaming-sarah.jpg) |
| Void Terrarium | Dead Cells |
| :------------------------------------------------------------------------: | :------------------------------------------------------------------: |
| ![Void Terrarium](./.github/images/void-terrarium.jpg) | ![Dead Cells](./.github/images/dead-cells.jpg) |
## Status
The emulator can currently load the `eboot.bin` of real games, execute native CPU instructions, and partially handle kernel-related functionality. However, several critical components are still missing.
@@ -59,33 +72,33 @@ Current capabilities include:
Some games have reached like `sceVideoOut` and AGC stages.
Currently the project primarily targets Windows. Cross-platform support (Linux and macOS) is planned, but development is currently focused on Windows to simplify early-stage debugging and iteration.
SharpEmu supports Windows, Linux, and macOS hosts. Video output uses Vulkan on
Windows and Linux, and MoltenVK on macOS. Platform support is still experimental,
so compatibility and performance vary by game, operating system, and GPU driver.
## Using
* Build or Publish project or download in release tab.
* Open Powershell.
* Run Emulator GUI.
* Or command: `.\SharpEmu "eboot.bin" 2>&1 | Tee-Object -FilePath "log.txt"`
Download the release archive for your operating system, extract it, and launch
SharpEmu with the path to a legally obtained game's `eboot.bin`.
## Games Tested
Windows PowerShell:
* **Demon's Souls Remake**
* [Demon's Souls [PPSA01341]](https://github.com/par274/sharpemu/issues/2)
* Demon's Souls is now video loop. Shaders are ready to be converted to SPIR-V/Vulkan. We are continuing our work on this.
![DeS videoOut submit first frame](./.github/images/des-videoout-shaders.jpg)
```powershell
.\SharpEmu.exe "C:\path\to\game\eboot.bin" 2>&1 |
Tee-Object -FilePath "SharpEmu.log"
```
* **Poppy Playtime Chapter 1**
* [Poppy Playtime Chapter 1 [PPSA20591]](https://github.com/par274/sharpemu/issues/3)
Linux and macOS:
* **SILENT HILL: The Short Message**
* [SILENT HILL: The Short Message [PPSA10112]](https://github.com/par274/sharpemu/issues/4)
```bash
chmod +x ./SharpEmu
* **Dreaming Sarah**
* [Dreaming Sarah [PPSA02929]](https://github.com/par274/sharpemu/issues/9)
* Real texture rendering for this game;
![Splash texture](./.github/images/dreaming-sarah.jpg)
./SharpEmu "/path/to/game/eboot.bin" 2>&1 |
tee SharpEmu.log
```
A Vulkan-capable GPU and current graphics driver are required. The macOS
release includes the MoltenVK Vulkan implementation.
> [!IMPORTANT]
> This project does **not** support or condone piracy.
@@ -94,8 +107,8 @@ Currently the project primarily targets Windows. Cross-platform support (Linux a
## Build
1. Install the **.NET SDK**.
2. Clone the repository: `git clone https://github.com/par274/sharpemu.git`
1. Install the .NET SDK version specified in [`global.json`](./global.json).
2. Clone the repository: `git clone https://github.com/sharpemu/sharpemu.git`
3. Open the solution file (`SharpEmu.slnx`) in **VSCode**.
4. Build the project: `dotnet build` or `dotnet publish`
5. Build artifacts will be located in the `artifacts` directory.
@@ -121,7 +134,7 @@ Provided valuable references for filesystem handling and low-level C# implementa
# License
- [**GPL-2.0 license**](https://github.com/par274/sharpemu/blob/main/LICENSE)
- [**GPL-2.0 license**](https://github.com/sharpemu/sharpemu/blob/main/LICENSE)
## Contributing
+3 -1
View File
@@ -7,10 +7,12 @@ path = [
"global.json",
"**/packages.lock.json",
"scripts/ps5_names.txt",
"src/SharpEmu.HLE/Aerolib/aerolib.bin",
"src/SharpEmu.GUI/Languages/**",
"src/SharpEmu.ShaderCompiler.Metal/Templates/**",
"tests/SharpEmu.ShaderCompiler.Metal.Tests/Goldens/**",
"_logs/**",
".github/images/**",
".github/pull_request_template.md",
"assets/images/**"
]
precedence = "aggregate"
+11
View File
@@ -7,9 +7,20 @@ SPDX-License-Identifier: GPL-2.0-or-later
<Folder Name="/src/">
<Project Path="src/SharpEmu.CLI/SharpEmu.CLI.csproj" />
<Project Path="src/SharpEmu.Core/SharpEmu.Core.csproj" />
<Project Path="src/SharpEmu.DebugClient/SharpEmu.DebugClient.csproj" />
<Project Path="src/SharpEmu.Debugger/SharpEmu.Debugger.csproj" />
<Project Path="src/SharpEmu.GUI/SharpEmu.GUI.csproj" />
<Project Path="src/SharpEmu.HLE/SharpEmu.HLE.csproj" />
<Project Path="src/SharpEmu.Libs/SharpEmu.Libs.csproj" />
<Project Path="src/SharpEmu.Logging/SharpEmu.Logging.csproj" />
<Project Path="src/SharpEmu.ShaderCompiler/SharpEmu.ShaderCompiler.csproj" />
<Project Path="src/SharpEmu.ShaderCompiler.Metal/SharpEmu.ShaderCompiler.Metal.csproj" />
<Project Path="src/SharpEmu.ShaderCompiler.Vulkan/SharpEmu.ShaderCompiler.Vulkan.csproj" />
<Project Path="src/SharpEmu.SourceGenerators/SharpEmu.SourceGenerators.csproj" />
</Folder>
<Folder Name="/tests/">
<Project Path="tests/SharpEmu.Libs.Tests/SharpEmu.Libs.Tests.csproj" />
<Project Path="tests/SharpEmu.ShaderCompiler.Metal.Tests/SharpEmu.ShaderCompiler.Metal.Tests.csproj" />
<Project Path="tests/SharpEmu.SourceGenerators.Tests/SharpEmu.SourceGenerators.Tests.csproj" />
</Folder>
</Solution>
Binary file not shown.

After

Width:  |  Height:  |  Size: 653 B

Binary file not shown.

After

Width:  |  Height:  |  Size: 698 B

Binary file not shown.

After

Width:  |  Height:  |  Size: 802 B

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.9 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.2 KiB

+20
View File
@@ -0,0 +1,20 @@
<!--
Copyright (C) 2026 SharpEmu Emulator Project
SPDX-License-Identifier: GPL-2.0-or-later
-->
# Aerolib Catalog
```bash
# NID to export name
python scripts/aerolib_catalog.py lookup Zxa0VhQVTsk
# Export name to NID
python scripts/aerolib_catalog.py lookup sceKernelWaitSema
# Search export names
python scripts/aerolib_catalog.py search VideoOut --limit 20
# Export all NID/name pairs to artifacts/aerolib.txt
python scripts/aerolib_catalog.py export
```
+75
View File
@@ -0,0 +1,75 @@
<!--
Copyright (C) 2026 SharpEmu Emulator Project
SPDX-License-Identifier: GPL-2.0-or-later
-->
# Bink 2 bridge
Demon's Souls plays Bink 2 (.bk2) files through a Bink implementation linked
directly into eboot.bin. It does not use libSceVideodec, therefore an HLE video
decoder cannot observe or replace those frames.
SharpEmu observes successful guest .bk2 opens and, when a Bink decoder is
available, presents its decoded BGRA frames at the normal guest-flip boundary.
This preserves the game's own timing and lets the host Vulkan presenter display
the movie without trying to execute the PS5-specific Bink GPU decode path.
The default path decodes by calling FFmpeg's own C API directly from managed
code (`src/SharpEmu.Libs/Bink/FfmpegNativeBinkFrameSource.cs`, via the
[FFmpeg.AutoGen](https://github.com/Ruslan-B/FFmpeg.AutoGen) P/Invoke
bindings) against a custom FFmpeg build
(`github.com/sharpemu/ffmpeg-core`, LGPL-2.1) that adds a Bink 2 decoder to
FFmpeg 7.1.2; see "Supplying the FFmpeg libraries" below for where those
libraries come from. No proprietary RAD SDK is needed to build or run
SharpEmu, and there is no C/C++ code of SharpEmu's own involved in decoding
-- SharpEmu.CLI.csproj only downloads a prebuilt release archive.
Set `SHARPEMU_BINK_MODE=guest` to leave decoding to the Bink implementation
statically linked into the game instead. Set `skip` only when explicitly
testing a title whose cinematics are optional.
Set SHARPEMU_BINK_MODE=dummy to retain the open and show a built-in,
non-decoded placeholder frame. This requires no SDK, but is a visual diagnostic
only; it does not decode the movie or alter its game logic.
SHARPEMU_BINK_MODE=native is equivalent to the default and mainly useful for
being explicit about it.
The experimental `SHARPEMU_BINK_MODE=ffmpeg` override is unrelated to the
default path above: instead of calling into FFmpeg in-process, it spawns a
standalone `ffmpeg` executable and reads raw frames from its stdout
(`src/SharpEmu.Libs/Bink/FfmpegBinkFrameSource.cs`). SharpEmu searches
`SHARPEMU_FFMPEG_PATH`, the executable directory, its `ffmpeg` subdirectory,
and then `PATH` (plus a couple of common Homebrew paths on macOS). That
`ffmpeg` build must contain a Bink 2 decoder itself; a stock FFmpeg build that
only recognizes the Bink container is not sufficient. Most users want the
default `native` mode instead, which always has Bink 2 support since it's
built against `ffmpeg-core` specifically.
## Supplying the FFmpeg libraries
`dotnet publish` fetches a prebuilt release of `github.com/sharpemu/ffmpeg-core`
(the tag is pinned in `SharpEmu.CLI.csproj`'s `FfmpegRuntimeTag`, matched to
the `FFmpeg.AutoGen` package version in `Directory.Packages.props` -- both
need to agree on the same FFmpeg ABI) and copies its dynamically linked
libraries into a `plugins` folder next to the published executable. No C
toolchain is required to build SharpEmu; publishing just downloads a zip.
`plugins` is a loose, unpacked folder rather than something embedded in the
single-file bundle, so the OS loader can resolve the libraries' own
inter-dependencies (`avcodec` depends on `avutil`, etc.) itself.
A plain `dotnet publish` with no `-r` still works: it defaults to the host
machine's own RID (see `Directory.Build.props`), so it fetches the matching
`ffmpeg-core` archive and populates `plugins` without any extra flags.
Passing an explicit `-r <rid>` (e.g. to cross-publish `linux-x64` from
Windows) still overrides that default normally.
To use a different set of FFmpeg libraries, drop them into the published
`plugins` folder yourself (matching FFmpeg's own file-naming and versioning
conventions, e.g. `avformat-61.dll` / `libavformat.so.61` / matching
`.dylib`) -- `FfmpegNativeBinkFrameSource` points `ffmpeg.RootPath` at that
folder and does not otherwise care where the files came from.
If the libraries are absent or fail to load, `FfmpegNativeBinkFrameSource.TryOpen`
degrades gracefully: SharpEmu logs one informational line ("Bink2 bridge
could not open movie ...") and leaves the guest's own rendering path
untouched, rather than crashing.
+177
View File
@@ -0,0 +1,177 @@
<!--
Copyright (C) 2026 SharpEmu Emulator Project
SPDX-License-Identifier: GPL-2.0-or-later
-->
# Live debug server
SharpEmu can expose a **live debug server** so an external process can inspect
and control a running guest over TCP. The server lives in the emulator; the
companion `SharpEmu.DebugClient` executable is one client, and the wire protocol
is simple enough to script against directly.
This document describes the moving parts and the wire protocol. For day-to-day
client usage, see
[`src/SharpEmu.DebugClient/DEVELOPER_READ.md`](../src/SharpEmu.DebugClient/DEVELOPER_READ.md).
## Layering
| Assembly | Role |
| -------- | ---- |
| `SharpEmu.Core` | Defines the dispatcher seam `ICpuDebugHook` / `ICpuDebugFrame` (namespace `SharpEmu.Core.Cpu.Debug`) and the `CpuExecutionOptions.DebugHook` slot. Core has **no** reference to the debugger. |
| `SharpEmu.Debugger` | The debugger: `DebuggerSession` (implements the hook), `BreakpointStore`, the TCP `DebuggerServer`, the pluggable `IDebugProtocol` with a JSON-lines implementation, and the `DebuggerServerHost` one-call wiring. |
| `SharpEmu.CLI` | Parses `--debug-server`, builds a `DebuggerServerHost`, hands its `Hook` to `SharpEmuRuntimeOptions.DebugHook`, and manages its lifetime. |
| `SharpEmu.DebugClient` | A standalone client executable. Depends only on the BCL. |
The dependency direction is important: Core stays debugger-agnostic and only
publishes the seam. Anything that observes execution implements
`ICpuDebugHook` and is injected through the options, so the debugger can evolve
without touching the CPU core.
## Execution model
`CpuDispatcher` enters a fresh frame for the process entry point and for each
module initializer. When a `DebugHook` is attached it is notified at those
boundaries:
- `OnFrameEnter(frame)` — before the native backend runs the frame. The
`DebuggerSession` decides whether to stop (pause request, breakpoint on the
entry address, single-step, or stop-at-entry). To stop, it **parks the
emulation thread** inside this call on a gate; the frame stays live, so a
client can read and write registers and memory while parked. `continue` /
`step` release the gate.
- `OnFrameExit(frame, result)` — after the frame completes.
Because pausing parks the one thread that owns the guest context, register and
memory accessors are only served while the session reports `Paused`; otherwise
they return "not paused" so a client never observes torn state.
### What is and isn't live yet
- **Live:** attach/handshake, run-state tracking, register read/write, memory
read/write, breakpoint management, execution breakpoints at frame entry,
pause, frame-level step, continue, and stop/resume/terminate events.
- **Surface only (armed as the backend grows hooks):** per-instruction
stepping and data watchpoints (`readwatch` / `writewatch` / `accesswatch`).
The verbs and types exist so clients and tooling can be written now.
## Enabling the server
```bash
SharpEmu --debug-server "/path/to/eboot.bin" # 127.0.0.1:5714
SharpEmu --debug-server=0.0.0.0:5714 "/path/to/eboot.bin"
```
The bind address defaults to loopback; a routable address must be given
explicitly. With stop-at-entry (the default `DebuggerSessionOptions.StopAtEntry`),
the guest parks at its first frame until a client connects and issues
`continue`, giving you a window to set breakpoints before any guest code runs.
## Browser frontend
The dependency-free Python frontend can choose and launch an `eboot.bin`, attach
to its debugger automatically, and provides execution controls, registers,
memory inspection, breakpoint management, process output, and a live protocol
activity stream:
```bash
./tools/SharpEmu.DebuggerFrontend/run.sh
```
It connects to `127.0.0.1:5714` and opens `http://127.0.0.1:8765/` by default.
See [`tools/SharpEmu.DebuggerFrontend/README.md`](../tools/SharpEmu.DebuggerFrontend/README.md)
for configuration and testing options.
## Wire protocol (json-lines/1)
One JSON object per line, UTF-8, `\n`-terminated, in both directions.
### Requests
A `command` string plus command-specific fields. Numeric fields accept a JSON
number or a `0x`-prefixed hex string.
| `command` | Fields | Reply `data` |
| --------- | ------ | ------------ |
| `ping` | — | — |
| `status` (`info`) | — | `state`, `breakpoints`, `lastStop?` |
| `state` | — | `state` |
| `registers` (`regs`) | — | `registers` (rax..r15, rip, rflags, fs_base, gs_base) |
| `set-register` | `register`, `value` | — |
| `read-memory` | `address`, `length` (≤ 65536) | `address`, `length`, `bytes` (hex) |
| `write-memory` | `address`, `bytes` (hex) | `written` |
| `list-breakpoints` (`breakpoints`) | — | `breakpoints[]` |
| `add-breakpoint` (`break`) | `address`, `kind?`, `length?` | `breakpoint` |
| `remove-breakpoint` (`delete-breakpoint`) | `id` | — |
| `enable-breakpoint` | `id`, `enabled?` (default true) | — |
| `continue` (`cont`, `c`) | — | — |
| `step` (`s`) | — | — |
| `pause` | — | — |
### Replies
```json
{"ok":true,"command":"registers","data":{ "registers": { "rax":"0x…", } }}
{"ok":false,"command":"read-memory","error":"Target is not paused."}
```
### Events (unsolicited)
```json
{"event":"hello","protocol":"json-lines/1","state":"Paused"}
{"event":"stopped","reason":"Breakpoint","address":"0x…","frameKind":"ProcessEntry","frameLabel":"eboot.bin","registers":{},"breakpoint":{}}
{"event":"resumed"}
{"event":"terminated"}
```
`reason` is one of `EntryPoint`, `Breakpoint`, `Watchpoint`, `Step`, `Pause`,
`Fault`, or `Stall`.
Stall stops include structured evidence in addition to the human-readable
detail. Import-loop evidence identifies the NID, resolved HLE export, repeating
guest return site, dispatch count, and first two ABI arguments:
```json
{
"event": "stopped",
"reason": "Stall",
"stall": {
"kind": "ImportLoop",
"nid": "9UK1vLZQft4",
"instructionPointer": "0x0000000801CE2418",
"dispatchIndex": 40667904,
"argument0": "0x0000000812345000",
"argument1": "0x0000000000000000",
"resolved": true,
"library": "libKernel",
"function": "scePthreadMutexLock"
}
}
```
The Python frontend uses this evidence to explain the likely failure class and
rank concrete checks/fixes. Its diagnosis is intentionally labelled heuristic:
it helps locate the responsible HLE/scheduler path but does not replace tracing.
## Swapping the protocol
`DebuggerServer` takes an `IDebugProtocol` factory. The default is
`JsonLineDebugProtocol`; a GDB remote serial stub (or any other framing) can be
dropped in without changing the session or command semantics, which live in
`DebugCommandDispatcher`.
## Embedding the server
```csharp
using SharpEmu.Debugger;
using SharpEmu.Core.Runtime;
await using var host = new DebuggerServerHost();
host.Start();
var options = new SharpEmuRuntimeOptions { DebugHook = host.Hook };
using var runtime = SharpEmuRuntime.CreateDefault(options);
var result = runtime.Run(ebootPath);
host.NotifyRunCompleted();
```
+57
View File
@@ -0,0 +1,57 @@
<!--
Copyright (C) 2026 SharpEmu Emulator Project
SPDX-License-Identifier: GPL-2.0-or-later
-->
# Guest write watch
`GuestWriteWatch` is an optional diagnostic tool. It helps you find managed
code and HLE code that damage guest memory. The tool starts only if you set one
or more `SHARPEMU_WATCH_*` environment variables.
The tool monitors writes through the SharpEmu managed virtual-memory APIs. It
does not monitor stores that native guest code makes directly. Use a platform
debugger or a hardware watchpoint to monitor these stores.
## Watch modes
- `SHARPEMU_WATCH_WRITE=0x<address>` logs a write that overlaps the eight-byte
block at the specified guest address.
- `SHARPEMU_WATCH_POOL_HEADER=1` monitors the pointer at offset `0x40`. It
monitors the first 64 direct mappings that have a size of 64 KiB and
protection value `0xF2`.
- `SHARPEMU_WATCH_VALUE_PATTERN=1` logs an eight-byte write if its lower 32 bits
are `1`. The upper 32 bits must look like a small guest-pointer prefix.
- `SHARPEMU_WATCH_VALUE1=1` logs short writes of value `1` in the high guest
memory range. The tool logs a maximum of 128 entries for each process.
- `SHARPEMU_WATCH_BULK_TORN=1` scans aligned 64-bit words in bulk writes. It
finds damaged pointer patterns and byte-shifted pointer patterns. The tool
logs a maximum of 64 entries for each process.
- `SHARPEMU_WATCH_BULK_DEST_HI=0x<high-dword>` scans only writes that have the
specified upper 32 bits in the destination address.
For each match, the tool logs the destination address, the data pattern, and the
managed call stack. The log uses the `watch_write` or `watch_bulk_torn` warning
tag.
Use these variables together to scan bulk writes in the
`0x00000080xxxxxxxx` region.
macOS and Linux:
```sh
SHARPEMU_WATCH_BULK_TORN=1 \
SHARPEMU_WATCH_BULK_DEST_HI=0x80 \
SharpEmu /path/to/eboot.bin
```
Windows PowerShell:
```powershell
$env:SHARPEMU_WATCH_BULK_TORN = "1"
$env:SHARPEMU_WATCH_BULK_DEST_HI = "0x80"
& .\SharpEmu.exe C:\path\to\game\eboot.bin
```
To reduce unnecessary log entries, use an exact `SHARPEMU_WATCH_WRITE`
address from a crash dump.
+90
View File
@@ -0,0 +1,90 @@
<!--
Copyright (C) 2026 SharpEmu Emulator Project
SPDX-License-Identifier: GPL-2.0-or-later
-->
## Release Script
SharpEmu releases are prepared and published using `scripts/release.py`.
The release process consists of two steps:
1. Prepare the version bump through a pull request.
2. Create and push the release tag after the pull request has been merged.
### Preparing a Release
Run:
```bash
python scripts/release.py prepare 0.0.2-beta.2
```
This command will:
- Verify that the working tree is clean.
- Update the local `main` branch.
- Create a new branch named `release/0.0.2-beta.2`.
- Update `SharpEmuVersion` in `Directory.Build.props`.
- Create a version bump commit.
- Push the release branch to the remote repository.
Afterwards, open a pull request from:
```text
release/0.0.2-beta.2
```
into:
```text
main
```
### Publishing a Release
Once the pull request has been merged, update your local repository:
```bash
git switch main
git pull --ff-only
```
Then create and push the release tag:
```bash
python scripts/release.py tag 0.0.2-beta.2
```
This command will:
- Verify that the working tree is clean.
- Confirm that `Directory.Build.props` contains the requested version.
- Create an annotated Git tag (`v0.0.2-beta.2`).
- Push the tag to the remote repository.
Pushing the tag automatically triggers the GitHub Release workflow.
### Version Format
Specify the version **without** the `v` prefix.
Examples:
```text
0.0.2
0.0.2-alpha.1
0.0.2-beta.1
0.0.2-beta.2
0.0.2-rc.1
```
The script automatically prefixes the Git tag with `v`.
### Notes
- Run `prepare` only from the `main` branch.
- Run `tag` only after the version bump pull request has been merged.
- Do not create release tags manually before merging the version bump.
- Both commands require a clean working tree.
- The version in `Directory.Build.props` must exactly match the version passed to the `tag` command.
+1 -1
View File
@@ -1,6 +1,6 @@
{
"sdk": {
"version": "10.0.103",
"rollForward": "disable"
"rollForward": "latestFeature"
}
}
+181
View File
@@ -0,0 +1,181 @@
#!/usr/bin/env python3
# Copyright (C) 2026 SharpEmu Emulator Project
# SPDX-License-Identifier: GPL-2.0-or-later
from __future__ import annotations
import argparse
import base64
import hashlib
import re
import sys
from pathlib import Path
NID_SUFFIX = bytes.fromhex("518d64a635ded8c1e6b039b1c3e55230")
NID_PATTERN = re.compile(r"^[A-Za-z0-9+-]{11}$")
DEFAULT_NAMES_FILE = Path(__file__).resolve().with_name("ps5_names.txt")
DEFAULT_EXPORT_FILE = Path(__file__).resolve().parents[1] / "artifacts" / "aerolib.txt"
def compute_nid(export_name: str) -> str:
digest = hashlib.sha1(export_name.encode("utf-8") + NID_SUFFIX).digest()
encoded = base64.b64encode(digest[:8][::-1]).decode("ascii")
return encoded.rstrip("=").replace("/", "-")
def read_names(path: Path) -> list[str]:
try:
return [
line.strip()
for line in path.read_text(encoding="utf-8").splitlines()
if line.strip()
]
except OSError as error:
raise SystemExit(f"Unable to read catalog '{path}': {error}") from error
def write_pair(nid: str, export_name: str) -> None:
print(f"{nid}\t{export_name}")
def lookup(args: argparse.Namespace) -> int:
value = args.value.strip()
if NID_PATTERN.fullmatch(value):
for export_name in read_names(args.names):
if compute_nid(export_name) == value:
write_pair(value, export_name)
return 0
print(f"NID not found in catalog: {value}", file=sys.stderr)
return 1
names = set(read_names(args.names))
write_pair(compute_nid(value), value)
if value not in names:
print("Warning: export name is not present in the catalog.", file=sys.stderr)
return 0
def search(args: argparse.Namespace) -> int:
names = read_names(args.names)
if args.regex:
try:
pattern = re.compile(args.query, 0 if args.case_sensitive else re.IGNORECASE)
except re.error as error:
print(f"Invalid regular expression: {error}", file=sys.stderr)
return 2
matches = (name for name in names if pattern.search(name))
elif args.case_sensitive:
matches = (name for name in names if args.query in name)
else:
query = args.query.casefold()
matches = (name for name in names if query in name.casefold())
count = 0
for export_name in matches:
write_pair(compute_nid(export_name), export_name)
count += 1
if args.limit and count >= args.limit:
break
if count == 0:
print(f"No catalog names matched: {args.query}", file=sys.stderr)
return 1
return 0
def export_catalog(args: argparse.Namespace) -> int:
pairs = [(compute_nid(name), name) for name in read_names(args.names)]
if args.sort == "nid":
pairs.sort(key=lambda pair: (pair[0], pair[1]))
elif args.sort == "name":
pairs.sort(key=lambda pair: pair[1])
args.output.parent.mkdir(parents=True, exist_ok=True)
try:
with args.output.open("w", encoding="utf-8", newline="\n") as output:
output.write("# NID\tExportName\n")
for nid, export_name in pairs:
output.write(f"{nid}\t{export_name}\n")
except OSError as error:
print(f"Unable to write catalog '{args.output}': {error}", file=sys.stderr)
return 1
print(f"Wrote {len(pairs)} entries to {args.output}")
return 0
def create_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(
description="Inspect the SharpEmu PS5 export-name/NID catalog.",
epilog=(
"Examples:\n"
" python scripts/aerolib_catalog.py lookup Zxa0VhQVTsk\n"
" python scripts/aerolib_catalog.py lookup sceKernelWaitSema\n"
" python scripts/aerolib_catalog.py search VideoOut --limit 20\n"
" python scripts/aerolib_catalog.py export"
),
formatter_class=argparse.RawDescriptionHelpFormatter,
)
parser.add_argument(
"--names",
type=Path,
default=DEFAULT_NAMES_FILE,
help=f"source name list (default: {DEFAULT_NAMES_FILE})",
)
subparsers = parser.add_subparsers(dest="command", required=True)
lookup_parser = subparsers.add_parser(
"lookup", help="resolve a NID or calculate the NID for an export name"
)
lookup_parser.add_argument("value", help="11-character NID or exact export name")
lookup_parser.set_defaults(handler=lookup)
search_parser = subparsers.add_parser(
"search", help="find export names and print matching NID/name pairs"
)
search_parser.add_argument("query", help="name substring or regular expression")
search_parser.add_argument(
"--limit", type=int, default=50, help="maximum matches; 0 means unlimited"
)
search_parser.add_argument(
"--case-sensitive", action="store_true", help="match case exactly"
)
search_parser.add_argument(
"--regex", action="store_true", help="treat the query as a regular expression"
)
search_parser.set_defaults(handler=search)
export_parser = subparsers.add_parser(
"export", help="write every NID/name pair to a tab-separated text file"
)
export_parser.add_argument(
"output",
type=Path,
nargs="?",
default=DEFAULT_EXPORT_FILE,
help=f"output file (default: {DEFAULT_EXPORT_FILE})",
)
export_parser.add_argument(
"--sort",
choices=("source", "nid", "name"),
default="nid",
help="output ordering (default: nid)",
)
export_parser.set_defaults(handler=export_catalog)
return parser
def main() -> int:
parser = create_parser()
args = parser.parse_args()
return args.handler(args)
if __name__ == "__main__":
raise SystemExit(main())
+37
View File
@@ -0,0 +1,37 @@
#!/usr/bin/env bash
# Copyright (C) 2026 SharpEmu Emulator Project
# SPDX-License-Identifier: GPL-2.0-or-later
#
# Downloads the official (universal x86_64+arm64) MoltenVK dylib and stages
# it next to a SharpEmu build as libvulkan.1.dylib. The macOS build runs as
# an x86-64 process under Rosetta 2, so Homebrew's arm64-only Vulkan
# libraries cannot be used; the presenter looks for this app-local copy.
#
# Usage: scripts/fetch-macos-moltenvk.sh [output-dir]
# (default output: artifacts/bin/Debug/net10.0/osx-x64)
set -euo pipefail
MVK_VERSION="${MVK_VERSION:-v1.4.0}"
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
OUT_DIR="${1:-$REPO_ROOT/artifacts/bin/Debug/net10.0/osx-x64}"
if [[ ! -d "$OUT_DIR" ]]; then
echo "output directory does not exist: $OUT_DIR (build first?)" >&2
exit 2
fi
WORK_DIR="$(mktemp -d)"
trap 'rm -rf "$WORK_DIR"' EXIT
echo ">> Downloading MoltenVK $MVK_VERSION..."
curl -sL -o "$WORK_DIR/mvk.tar" \
"https://github.com/KhronosGroup/MoltenVK/releases/download/$MVK_VERSION/MoltenVK-macos.tar"
tar -xf "$WORK_DIR/mvk.tar" -C "$WORK_DIR" \
MoltenVK/MoltenVK/dynamic/dylib/macOS/libMoltenVK.dylib
DYLIB="$WORK_DIR/MoltenVK/MoltenVK/dynamic/dylib/macOS/libMoltenVK.dylib"
file "$DYLIB" | grep -q x86_64 || { echo "downloaded dylib lacks x86_64 slice" >&2; exit 3; }
cp "$DYLIB" "$OUT_DIR/libMoltenVK.dylib"
cp "$DYLIB" "$OUT_DIR/libvulkan.1.dylib"
echo ">> Staged libMoltenVK.dylib + libvulkan.1.dylib in $OUT_DIR"
-54
View File
@@ -1,54 +0,0 @@
# Copyright (C) 2026 SharpEmu Emulator Project
# SPDX-License-Identifier: GPL-2.0-or-later
#!/usr/bin/env python3
import hashlib
import struct
from base64 import b64encode as base64enc
from binascii import unhexlify as uhx
from pathlib import Path
NAMES = 'scripts/ps5_names.txt'
OUTPUT = 'src/SharpEmu.HLE/Aerolib/aerolib.bin'
def name2nid(name):
symbol = hashlib.sha1(name.encode() + uhx('518D64A635DED8C1E6B039B1C3E55230')).digest()
id_val = struct.unpack('<Q', symbol[:8])[0]
nid = base64enc(uhx('%016x' % id_val), b'+-').rstrip(b'=')
return nid.decode('utf-8')
def generate():
names_path = Path(NAMES)
output_path = Path(OUTPUT)
entries = []
with open(names_path, 'r', encoding='utf-8') as f:
for line in f:
name = line.strip()
if name:
nid = name2nid(name)
entries.append((nid, name))
print(f"Found {len(entries)} entries")
data = bytearray()
data.extend(struct.pack('<I', len(entries)))
for nid, name in entries:
nid_bytes = nid.encode('utf-8')
name_bytes = name.encode('utf-8')
data.append(len(nid_bytes))
data.extend(nid_bytes)
data.extend(struct.pack('<H', len(name_bytes)))
data.extend(name_bytes)
output_path.parent.mkdir(parents=True, exist_ok=True)
with open(output_path, 'wb') as f:
f.write(data)
print(f"Generated: {output_path} ({len(data):,} bytes)")
print(f"Total entries: {len(entries)}")
if __name__ == "__main__":
generate()
+6
View File
@@ -154449,3 +154449,9 @@ vector_str_substr
WTFAnnotateBenignRaceSized
WTFAnnotateHappensAfter
WTFAnnotateHappensBefore
_sceUlobjmgrRegisterObject
_sceUlobjmgrUnregisterObject
sceNpEAAccessInitialize
sceNpEAAccessTerminate
sceNpHasEAAccessSubscription
sceNpHasEAAccessSubscriptionAbortRequest
+408
View File
@@ -0,0 +1,408 @@
#!/usr/bin/env python3
# Copyright (C) 2026 SharpEmu Emulator Project
# SPDX-License-Identifier: GPL-2.0-or-later
from __future__ import annotations
import argparse
import re
import subprocess
import sys
from pathlib import Path
VERSION_PATTERN = re.compile(r"^\d+\.\d+\.\d+(?:-[0-9A-Za-z.-]+)?$")
VERSION_ELEMENT_PATTERN = re.compile(
r"(<SharpEmuVersion>)([^<]+)(</SharpEmuVersion>)"
)
class ReleaseError(RuntimeError):
pass
def run_git(
*args: str,
cwd: Path,
capture_output: bool = False,
) -> str:
command = ["git", *args]
try:
result = subprocess.run(
command,
cwd=cwd,
check=True,
text=True,
capture_output=capture_output,
)
except FileNotFoundError:
raise ReleaseError("Git was not found in PATH.") from None
except subprocess.CalledProcessError as error:
stderr = error.stderr.strip() if error.stderr else ""
detail = f"\n{stderr}" if stderr else ""
raise ReleaseError(
f"Git command failed: {' '.join(command)}{detail}"
) from error
return result.stdout.strip() if capture_output else ""
def find_repository_root(script_path: Path) -> Path:
root = run_git(
"rev-parse",
"--show-toplevel",
cwd=script_path.resolve().parent,
capture_output=True,
)
return Path(root)
def get_status(repository_root: Path) -> str:
return run_git(
"status",
"--porcelain",
cwd=repository_root,
capture_output=True,
)
def ensure_clean_worktree(repository_root: Path) -> None:
status = get_status(repository_root)
if status:
raise ReleaseError(
"The working tree is not clean.\n\n"
f"{status}\n\n"
"Commit, stash, or remove these changes first."
)
def get_current_branch(repository_root: Path) -> str:
branch = run_git(
"branch",
"--show-current",
cwd=repository_root,
capture_output=True,
)
if not branch:
raise ReleaseError(
"HEAD is detached. Switch to a branch before continuing."
)
return branch
def ensure_branch_does_not_exist(
repository_root: Path,
branch: str,
remote: str,
) -> None:
local_branch = run_git(
"branch",
"--list",
branch,
cwd=repository_root,
capture_output=True,
)
if local_branch:
raise ReleaseError(f"Branch {branch} already exists locally.")
remote_branch = run_git(
"ls-remote",
"--heads",
remote,
branch,
cwd=repository_root,
capture_output=True,
)
if remote_branch:
raise ReleaseError(
f"Branch {branch} already exists on {remote}."
)
def ensure_tag_does_not_exist(
repository_root: Path,
tag: str,
remote: str,
) -> None:
local_tag = run_git(
"tag",
"--list",
tag,
cwd=repository_root,
capture_output=True,
)
if local_tag:
raise ReleaseError(f"Tag {tag} already exists locally.")
remote_tag = run_git(
"ls-remote",
"--tags",
remote,
f"refs/tags/{tag}",
cwd=repository_root,
capture_output=True,
)
if remote_tag:
raise ReleaseError(f"Tag {tag} already exists on {remote}.")
def read_version(props_path: Path) -> str:
if not props_path.exists():
raise ReleaseError(f"Version file not found: {props_path}")
content = props_path.read_text(encoding="utf-8")
match = VERSION_ELEMENT_PATTERN.search(content)
if match is None:
raise ReleaseError(
f"SharpEmuVersion was not found in {props_path.name}."
)
return match.group(2).strip()
def update_version(props_path: Path, version: str) -> str:
content = props_path.read_text(encoding="utf-8")
current_version = read_version(props_path)
if current_version == version:
raise ReleaseError(
f"SharpEmuVersion is already set to {version}."
)
updated_content, replacement_count = VERSION_ELEMENT_PATTERN.subn(
rf"\g<1>{version}\g<3>",
content,
count=1,
)
if replacement_count != 1:
raise ReleaseError(
"Expected exactly one SharpEmuVersion element."
)
props_path.write_text(
updated_content,
encoding="utf-8",
newline="\n",
)
return current_version
def prepare_release(
repository_root: Path,
props_path: Path,
version: str,
remote: str,
) -> None:
ensure_clean_worktree(repository_root)
current_branch = get_current_branch(repository_root)
if current_branch != "main":
raise ReleaseError(
f"Prepare must be run from main, not {current_branch}."
)
run_git(
"pull",
"--ff-only",
remote,
"main",
cwd=repository_root,
)
branch = f"release/{version}"
ensure_branch_does_not_exist(
repository_root,
branch,
remote,
)
previous_version = read_version(props_path)
run_git(
"switch",
"-c",
branch,
cwd=repository_root,
)
try:
update_version(props_path, version)
relative_props_path = props_path.relative_to(repository_root)
run_git(
"add",
relative_props_path.as_posix(),
cwd=repository_root,
)
run_git(
"commit",
"-m",
f"chore: bump version to {version}",
cwd=repository_root,
)
run_git(
"push",
"-u",
remote,
branch,
cwd=repository_root,
)
except Exception:
print(
"\nPrepare failed. The release branch may still exist locally.",
file=sys.stderr,
)
raise
print()
print(f"Prepared release {previous_version} -> {version}")
print(f"Branch pushed: {branch}")
print()
print("Open a pull request from:")
print(f" {branch}")
print("into:")
print(" main")
print()
print("After merging the PR, run:")
print(f" python scripts/release.py tag {version}")
def create_release_tag(
repository_root: Path,
props_path: Path,
version: str,
remote: str,
) -> None:
ensure_clean_worktree(repository_root)
current_branch = get_current_branch(repository_root)
if current_branch != "main":
raise ReleaseError(
f"Tagging must be run from main, not {current_branch}."
)
run_git(
"pull",
"--ff-only",
remote,
"main",
cwd=repository_root,
)
current_version = read_version(props_path)
if current_version != version:
raise ReleaseError(
"Version mismatch:\n"
f" Directory.Build.props: {current_version}\n"
f" Requested tag: {version}"
)
tag = f"v{version}"
ensure_tag_does_not_exist(
repository_root,
tag,
remote,
)
run_git(
"tag",
"-a",
tag,
"-m",
f"SharpEmu {version}",
cwd=repository_root,
)
run_git(
"push",
remote,
tag,
cwd=repository_root,
)
print()
print(f"Successfully pushed tag {tag}.")
print("The release workflow should start automatically.")
def parse_arguments() -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Prepare or tag a SharpEmu release."
)
subparsers = parser.add_subparsers(
dest="command",
required=True,
)
for command in ("prepare", "tag"):
subparser = subparsers.add_parser(command)
subparser.add_argument(
"version",
help="Version without the v prefix, e.g. 0.0.2-beta.2.",
)
subparser.add_argument(
"--remote",
default="origin",
help="Git remote. Default: origin.",
)
arguments = parser.parse_args()
if not VERSION_PATTERN.fullmatch(arguments.version):
parser.error(
"Version must look like 0.0.2, "
"0.0.2-beta.2, or 0.0.2-rc.1."
)
return arguments
def main() -> int:
arguments = parse_arguments()
try:
repository_root = find_repository_root(Path(__file__))
props_path = repository_root / "Directory.Build.props"
if arguments.command == "prepare":
prepare_release(
repository_root,
props_path,
arguments.version,
arguments.remote,
)
else:
create_release_tag(
repository_root,
props_path,
arguments.version,
arguments.remote,
)
except ReleaseError as error:
print(f"Error: {error}", file=sys.stderr)
return 1
return 0
if __name__ == "__main__":
raise SystemExit(main())
+32
View File
@@ -0,0 +1,32 @@
#!/usr/bin/env bash
# Copyright (C) 2026 SharpEmu Emulator Project
# SPDX-License-Identifier: GPL-2.0-or-later
#
# Smoke-tests the linux-x64 build inside an amd64 container. Useful from any
# host (including Apple Silicon, where Docker runs the amd64 image under
# emulation) to confirm the cross-platform layer keeps working on Linux.
#
# Usage: scripts/test-linux-docker.sh /path/to/eboot.bin
set -euo pipefail
GAME_PATH="${1:-}"
if [[ -z "$GAME_PATH" || ! -f "$GAME_PATH" ]]; then
echo "usage: $0 <path-to-eboot.bin>" >&2
exit 2
fi
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
GAME_DIR="$(cd "$(dirname "$GAME_PATH")" && pwd)"
GAME_FILE="$(basename "$GAME_PATH")"
PUBLISH_DIR="$REPO_ROOT/artifacts/publish/SharpEmu.CLI/Debug/net10.0/linux-x64"
echo ">> Publishing linux-x64 self-contained build..."
dotnet publish "$REPO_ROOT/src/SharpEmu.CLI" \
-c Debug -r linux-x64 --self-contained -p:PublishSingleFile=false
echo ">> Running inside linux/amd64 container..."
docker run --rm --platform linux/amd64 \
-v "$PUBLISH_DIR":/app:ro \
-v "$GAME_DIR":/game:ro \
mcr.microsoft.com/dotnet/runtime-deps:10.0 \
/app/SharpEmu --log-level=info "/game/$GAME_FILE"
+55
View File
@@ -0,0 +1,55 @@
#!/usr/bin/env bash
# Copyright (C) 2026 SharpEmu Emulator Project
# SPDX-License-Identifier: GPL-2.0-or-later
set -euo pipefail
if [ "$#" -ne 4 ]; then
echo "usage: $0 <spirv-val> <expected-version> <target-env> <module-directory>" >&2
exit 2
fi
validator=$1
expected_version=$2
target_env=$3
module_directory=$4
if [ ! -x "$validator" ]; then
echo "SPIR-V validator is not executable: $validator" >&2
exit 2
fi
if [ ! -d "$module_directory" ]; then
echo "SPIR-V module directory does not exist: $module_directory" >&2
exit 2
fi
validator_version="$("$validator" --version | head -n 1)"
if [[ "$validator_version" != *"SPIRV-Tools $expected_version"* ]]; then
echo "unexpected SPIRV-Tools version: $validator_version (expected $expected_version)" >&2
exit 2
fi
echo "Validator: $validator_version"
echo "Target environment: $target_env"
mapfile -d '' modules < <(find "$module_directory" -type f -name '*.spv' -print0 | sort -z)
if [ "${#modules[@]}" -eq 0 ]; then
echo "no SPIR-V modules found in $module_directory" >&2
exit 1
fi
failures=0
for module in "${modules[@]}"; do
echo "Validating module: $module"
if ! "$validator" --target-env "$target_env" "$module"; then
echo "SPIR-V validation failed: $module" >&2
failures=1
fi
done
if [ "$failures" -ne 0 ]; then
exit 1
fi
echo "Validated ${#modules[@]} synthetic SPIR-V modules."
+376 -54
View File
@@ -5,6 +5,7 @@ using SharpEmu.Core.Runtime;
using SharpEmu.Core.Cpu;
using SharpEmu.GUI;
using SharpEmu.HLE;
using SharpEmu.Libs.VideoOut;
using SharpEmu.Logging;
using System.Runtime.InteropServices;
using System.Text;
@@ -57,11 +58,16 @@ internal static partial class Program
private static int Run(string[] args)
{
args = NormalizeInternalArguments(args, out var isMitigatedChild);
if (args.Length == 0 && !isMitigatedChild)
if (Updater.TryApply(args, out var updateExitCode))
{
return updateExitCode;
}
args = NormalizeInternalArguments(args, out var isMitigatedChild);
PreloadGlfw();
if (args.Length == 0)
{
// No arguments: open the desktop frontend. Any argument selects
// the classic CLI behavior below.
return GuiLauncher.Run();
}
@@ -74,6 +80,163 @@ internal static partial class Program
TryEnableConsoleFileMirror(earlyLogFilePath);
}
if (!CheckHostArchitecture())
{
return 5;
}
if (OperatingSystem.IsMacOS() || OperatingSystem.IsLinux())
{
if (OperatingSystem.IsMacOS())
{
ConfigureMoltenVkDefaults();
PreloadMacVulkanLoader();
}
// GLFW requires window creation and event processing on the
// process main thread: AppKit demands it on macOS, and X11 has a
// single event queue that must be serviced from the main thread
// (a window created and polled off it may never map, which showed
// as a running game with no visible window on Linux). Emulation
// moves to a worker thread and the main thread services the window
// work the video presenter posts. Windows keeps a per-thread event
// queue, so its window stays on the presenter's own thread.
var exitCode = 0;
HostMainThread.Enable();
var emulation = new Thread(() =>
{
try
{
exitCode = RunEmulator(args, isMitigatedChild);
}
finally
{
HostMainThread.Shutdown();
}
}, 32 * 1024 * 1024)
{
Name = "SharpEmu Emulation",
};
emulation.Start();
HostMainThread.Pump();
emulation.Join();
return exitCode;
}
return RunEmulator(args, isMitigatedChild);
}
/// <summary>
/// The supported host execution model, checked before any emulation
/// starts: the CPU backend executes guest x86-64 code natively, so the
/// host process must be x86-64 — win-x64/linux-x64 on x64 hardware, or
/// osx-x64 under Rosetta 2 on Apple Silicon (Rosetta translates the
/// whole process, so it still reports as X64 here). An arm64 process
/// (e.g. the osx-arm64 build) can browse the GUI but cannot run games;
/// failing up front distinguishes that from MoltenVK, signal-handler,
/// or guest-memory startup problems.
/// </summary>
private static bool CheckHostArchitecture()
{
if (RuntimeInformation.ProcessArchitecture == Architecture.X64)
{
return true;
}
Console.Error.WriteLine(
$"[LOADER][ERROR] Unsupported process architecture " +
$"{RuntimeInformation.ProcessArchitecture}: guest code executes " +
"natively, so SharpEmu must run as an x86-64 process.");
if (OperatingSystem.IsMacOS())
{
Console.Error.WriteLine(
"[LOADER][ERROR] On Apple Silicon, use the osx-x64 build under " +
"Rosetta 2 (install with: softwareupdate --install-rosetta).");
}
return false;
}
/// <summary>
/// Applies MoltenVK performance defaults before the Vulkan loader is
/// loaded. Existing user-provided values always take precedence.
/// </summary>
private static void ConfigureMoltenVkDefaults()
{
try
{
_ = MacSetEnv("MVK_CONFIG_SYNCHRONOUS_QUEUE_SUBMITS", "0", 0);
_ = MacSetEnv("MVK_CONFIG_SHOULD_MAXIMIZE_CONCURRENT_COMPILATION", "1", 0);
_ = MacSetEnv("MVK_CONFIG_USE_METAL_ARGUMENT_BUFFERS", "1", 0);
_ = MacSetEnv("MVK_CONFIG_RESUME_LOST_DEVICE", "1", 0);
}
catch (Exception exception)
{
Console.Error.WriteLine(
$"[LOADER][WARN] Failed to set MoltenVK defaults: {exception.Message}");
}
}
/// <summary>
/// Makes a Vulkan loader visible to GLFW's dlopen("libvulkan.1.dylib").
/// Homebrew's Vulkan libraries are arm64-only and cannot load into this
/// x86-64 (Rosetta 2) process, so a universal libMoltenVK.dylib placed
/// next to the executable (named libvulkan.1.dylib) is preloaded here;
/// dyld then resolves GLFW's bare-name dlopen to the loaded image.
/// </summary>
private static void PreloadMacVulkanLoader()
{
var candidates = new[]
{
Path.Combine(AppContext.BaseDirectory, "libvulkan.1.dylib"),
Path.Combine(AppContext.BaseDirectory, "libMoltenVK.dylib"),
Path.Combine(
Environment.GetFolderPath(Environment.SpecialFolder.UserProfile),
".sharpemu", "x64lib", "libvulkan.1.dylib"),
};
foreach (var candidate in candidates)
{
if (File.Exists(candidate) && NativeLibrary.TryLoad(candidate, out _))
{
Console.Error.WriteLine($"[LOADER][INFO] Vulkan loader preloaded: {candidate}");
return;
}
}
if (NativeLibrary.TryLoad("libvulkan.1.dylib", out _))
{
return;
}
Console.Error.WriteLine(
"[LOADER][WARN] No x86-64 Vulkan loader found; video output will be unavailable. " +
"Place a universal libMoltenVK.dylib (from the MoltenVK releases) next to SharpEmu " +
"as libvulkan.1.dylib.");
}
/// <summary>
/// SharpEmu.CLI.csproj publishes glfw into a "plugins" subfolder rather
/// than flat next to the executable, which falls outside the default OS
/// DLL/dlopen search path. Preloading it here by full path first means
/// any later bare-name lookup (however Silk.NET/GLFW itself resolves the
/// library) finds it already loaded in the process and reuses it -- the
/// same technique <see cref="PreloadMacVulkanLoader"/> already relies on
/// for the Vulkan loader.
/// </summary>
private static void PreloadGlfw()
{
var fileName = OperatingSystem.IsWindows() ? "glfw3.dll"
: OperatingSystem.IsMacOS() ? "libglfw.3.dylib"
: "libglfw.so.3";
var candidate = Path.Combine(AppContext.BaseDirectory, "plugins", fileName);
if (File.Exists(candidate))
{
NativeLibrary.TryLoad(candidate, out _);
}
}
private static int RunEmulator(string[] args, bool isMitigatedChild)
{
Console.Error.WriteLine($"[DEBUG] SharpEmu starting with {args.Length} args");
if (!isMitigatedChild && TryRunMitigatedChild(args, out var childExitCode))
@@ -81,7 +244,17 @@ internal static partial class Program
return childExitCode;
}
if (!TryParseArguments(args, out var ebootPath, out var runtimeOptions, out var logLevel, out var logFilePath))
if (!TryExtractHostSurfaceArgument(args, out var emulatorArgs, out var hostSurface, out var hostSurfaceError))
{
Console.Error.WriteLine($"[LOADER][ERROR] {hostSurfaceError}");
return 1;
}
HostSessionControl.SetEmbeddedHostSurface(
hostSurface?.WindowHandle ?? 0,
hostSurface?.DisplayHandle ?? 0);
if (!TryParseArguments(emulatorArgs, out var ebootPath, out var runtimeOptions, out var logLevel, out var logFilePath))
{
PrintUsage();
return 1;
@@ -106,53 +279,156 @@ internal static partial class Program
return 2;
}
if (!TryGetDebugServerOptions(args, out var debugServerEnabled, out var debugServerOptions, out var debugServerError))
{
Log.Error($"Invalid --debug-server endpoint: {debugServerError}");
return 1;
}
SharpEmu.Debugger.DebuggerServerHost? debugHost = null;
if (debugServerEnabled)
{
debugHost = new SharpEmu.Debugger.DebuggerServerHost(debugServerOptions);
try
{
debugHost.Start();
Log.Info($"Live debug server listening on {debugHost.Endpoint}. Attach with SharpEmu.DebugClient.");
// With StopAtEntry, the guest parks at its first frame until a
// client connects and continues.
runtimeOptions = runtimeOptions with { DebugHook = debugHost.Hook };
}
catch (Exception ex)
{
Log.Error("Failed to start the debug server.", ex);
debugHost.DisposeAsync().AsTask().GetAwaiter().GetResult();
return 6;
}
}
Console.Error.WriteLine("[DEBUG] Creating runtime...");
using var runtime = SharpEmuRuntime.CreateDefault(runtimeOptions);
OrbisGen2Result result;
try
{
Console.Error.WriteLine($"[DEBUG] Running: {ebootPath}");
result = runtime.Run(ebootPath);
Console.Error.WriteLine($"[DEBUG] Result: {result}");
if (hostSurface is not null && !VulkanVideoHost.TryAttachSurface(hostSurface))
{
Console.Error.WriteLine("[LOADER][ERROR] The requested GUI host surface is already active.");
return 3;
}
using var runtime = SharpEmuRuntime.CreateDefault(runtimeOptions);
OrbisGen2Result result;
ConsoleCancelEventHandler? cancelHandler = null;
try
{
cancelHandler = (_, eventArgs) =>
{
eventArgs.Cancel = true;
VideoOutExports.NotifyHostInterrupt();
};
Console.CancelKeyPress += cancelHandler;
Console.Error.WriteLine($"[DEBUG] Running: {ebootPath}");
result = runtime.Run(ebootPath);
Console.Error.WriteLine($"[DEBUG] Result: {result}");
}
catch (Exception ex)
{
Console.Error.WriteLine($"[DEBUG] Exception: {ex}");
Log.Error("SharpEmu failed to run.", ex);
return 3;
}
finally
{
if (cancelHandler is not null)
{
Console.CancelKeyPress -= cancelHandler;
}
}
Log.Info($"SharpEmu execution completed. Result={result} (0x{(int)result:X8})");
if (!string.IsNullOrWhiteSpace(runtime.LastSessionSummary))
{
Log.Info(runtime.LastSessionSummary);
}
if (!string.IsNullOrWhiteSpace(runtime.LastBasicBlockTrace))
{
Log.Info("BB trace:");
Log.Info(runtime.LastBasicBlockTrace);
}
if (!string.IsNullOrWhiteSpace(runtime.LastMilestoneLog))
{
Log.Info(runtime.LastMilestoneLog);
}
if (result != OrbisGen2Result.ORBIS_GEN2_OK && !string.IsNullOrWhiteSpace(runtime.LastExecutionDiagnostics))
{
Log.Warn(runtime.LastExecutionDiagnostics);
}
if (runtimeOptions.ImportTraceLimit > 0 && !string.IsNullOrWhiteSpace(runtime.LastExecutionTrace))
{
Log.Info("Import trace:");
Log.Info(runtime.LastExecutionTrace);
}
return result == OrbisGen2Result.ORBIS_GEN2_OK ? 0 : 4;
}
catch (Exception ex)
finally
{
Console.Error.WriteLine($"[DEBUG] Exception: {ex}");
Log.Error("SharpEmu failed to run.", ex);
return 3;
if (debugHost is not null)
{
debugHost.NotifyRunCompleted();
debugHost.DisposeAsync().AsTask().GetAwaiter().GetResult();
}
HostSessionControl.SetEmbeddedHostSurface(0);
if (hostSurface is not null)
{
VulkanVideoHost.RequestClose();
VulkanVideoHost.DetachSurface(hostSurface);
hostSurface.Dispose();
}
}
}
private static bool TryExtractHostSurfaceArgument(
IReadOnlyList<string> args,
out string[] emulatorArgs,
out VulkanHostSurface? hostSurface,
out string? error)
{
const string hostSurfacePrefix = "--host-surface=";
var remaining = new List<string>(args.Count);
hostSurface = null;
error = null;
foreach (var argument in args)
{
if (!argument.StartsWith(hostSurfacePrefix, StringComparison.OrdinalIgnoreCase))
{
remaining.Add(argument);
continue;
}
if (hostSurface is not null)
{
emulatorArgs = [];
error = "more than one GUI host surface was specified";
return false;
}
var descriptor = argument[hostSurfacePrefix.Length..];
if (!VulkanHostSurface.TryCreateChildProcessSurface(descriptor, out hostSurface, out error))
{
emulatorArgs = [];
return false;
}
}
Log.Info($"SharpEmu execution completed. Result={result} (0x{(int)result:X8})");
if (!string.IsNullOrWhiteSpace(runtime.LastSessionSummary))
{
Log.Info(runtime.LastSessionSummary);
}
if (!string.IsNullOrWhiteSpace(runtime.LastBasicBlockTrace))
{
Log.Info("BB trace:");
Log.Info(runtime.LastBasicBlockTrace);
}
if (!string.IsNullOrWhiteSpace(runtime.LastMilestoneLog))
{
Log.Info(runtime.LastMilestoneLog);
}
if (result != OrbisGen2Result.ORBIS_GEN2_OK && !string.IsNullOrWhiteSpace(runtime.LastExecutionDiagnostics))
{
Log.Warn(runtime.LastExecutionDiagnostics);
}
if (runtimeOptions.ImportTraceLimit > 0 && !string.IsNullOrWhiteSpace(runtime.LastExecutionTrace))
{
Log.Info("Import trace:");
Log.Info(runtime.LastExecutionTrace);
}
return result == OrbisGen2Result.ORBIS_GEN2_OK ? 0 : 4;
emulatorArgs = remaining.ToArray();
return true;
}
private static void EnsureCliConsole()
@@ -249,7 +525,9 @@ internal static partial class Program
return handle != 0 && handle != -1;
}
private static string[] NormalizeInternalArguments(string[] args, out bool isMitigatedChild)
private static string[] NormalizeInternalArguments(
string[] args,
out bool isMitigatedChild)
{
isMitigatedChild = false;
var trustedMitigatedChild = string.Equals(
@@ -295,12 +573,7 @@ internal static partial class Program
return false;
}
var childArgs = new string[args.Length + 1];
childArgs[0] = MitigatedChildFlag;
for (var i = 0; i < args.Length; i++)
{
childArgs[i + 1] = args[i];
}
string[] childArgs = [MitigatedChildFlag, .. args];
var commandLine = BuildCommandLine(processPath, childArgs);
var startupInfoEx = new STARTUPINFOEX();
@@ -351,7 +624,7 @@ internal static partial class Program
nint jobHandle = 0;
Environment.SetEnvironmentVariable(MitigatedChildEnvironment, "1");
var created = CreateProcessW(
processPath,
null,
cmdLineBuilder,
0,
0,
@@ -769,8 +1042,45 @@ internal static partial class Program
private static void PrintUsage()
{
Log.Info("Usage: SharpEmu.CLI [--strict] [--trace-imports[=N]] [--cpu-engine=<native>] [--log-level=<level>] [--log-file[=<path>]] <path-to-eboot.bin>");
Log.Info("Usage: SharpEmu.CLI [--strict] [--trace-imports[=N]] [--cpu-engine=<native>] [--log-level=<level>] [--log-file[=<path>]] [--debug-server[=host:port]] <path-to-eboot.bin>");
Log.Info(@"Example: SharpEmu.CLI --cpu-engine=native --trace-imports=64 --log-level=debug --log-file ""E:\Games\...\eboot.bin""");
Log.Info("Debug server: --debug-server starts a live debug listener (default 127.0.0.1:5714); connect with SharpEmu.DebugClient.");
}
/// <summary>
/// Detects the <c>--debug-server</c> flag and parses its optional
/// <c>host:port</c> endpoint. Returns false only when the flag is present but
/// its endpoint is malformed, so the caller can abort with a clear error.
/// </summary>
private static bool TryGetDebugServerOptions(
string[] args,
out bool enabled,
out SharpEmu.Debugger.Server.DebuggerServerOptions options,
out string error)
{
enabled = false;
options = new SharpEmu.Debugger.Server.DebuggerServerOptions();
error = string.Empty;
foreach (var argument in args)
{
if (string.Equals(argument, "--debug-server", StringComparison.OrdinalIgnoreCase))
{
enabled = true;
continue;
}
const string prefix = "--debug-server=";
if (argument.StartsWith(prefix, StringComparison.OrdinalIgnoreCase))
{
enabled = true;
if (!SharpEmu.Debugger.Server.DebuggerServerOptions.TryParseEndpoint(argument[prefix.Length..], out options, out error))
{
return false;
}
}
}
return true;
}
private static bool TryParseArguments(
@@ -804,6 +1114,15 @@ internal static partial class Program
continue;
}
// The debug-server endpoint is parsed separately (see
// TryGetDebugServerOptions); accept the flag here so it is not
// rejected as an unknown option or mistaken for the eboot path.
if (string.Equals(argument, "--debug-server", StringComparison.OrdinalIgnoreCase) ||
argument.StartsWith("--debug-server=", StringComparison.OrdinalIgnoreCase))
{
continue;
}
if (string.Equals(argument, "--trace-imports", StringComparison.OrdinalIgnoreCase))
{
importTraceLimit = DefaultImportTraceLimit;
@@ -1131,7 +1450,7 @@ internal static partial class Program
[DllImport("kernel32.dll", EntryPoint = "CreateProcessW", SetLastError = true, CharSet = CharSet.Unicode)]
[return: MarshalAs(UnmanagedType.Bool)]
private static extern bool CreateProcessW(
string applicationName,
string? applicationName,
StringBuilder commandLine,
nint processAttributes,
nint threadAttributes,
@@ -1188,4 +1507,7 @@ internal static partial class Program
uint creationDisposition,
uint flagsAndAttributes,
nint templateFile);
[DllImport("libSystem", EntryPoint = "setenv")]
private static extern int MacSetEnv(string name, string value, int overwrite);
}
+86 -7
View File
@@ -7,6 +7,7 @@ SPDX-License-Identifier: GPL-2.0-or-later
<ItemGroup>
<ProjectReference Include="..\SharpEmu.Core\SharpEmu.Core.csproj" />
<ProjectReference Include="..\SharpEmu.Debugger\SharpEmu.Debugger.csproj" />
<ProjectReference Include="..\SharpEmu.GUI\SharpEmu.GUI.csproj" />
<ProjectReference Include="..\SharpEmu.Logging\SharpEmu.Logging.csproj" />
</ItemGroup>
@@ -16,35 +17,55 @@ SPDX-License-Identifier: GPL-2.0-or-later
console window; CLI mode re-attaches to the parent terminal's console. -->
<OutputType>WinExe</OutputType>
<AssemblyName>SharpEmu</AssemblyName>
<RuntimeIdentifiers>win-x64;linux-x64;osx-arm64</RuntimeIdentifiers>
<!-- osx-x64 is the macOS target: the CPU backend executes guest x86-64
natively, so on Apple Silicon it runs under Rosetta 2. -->
<RuntimeIdentifiers>win-x64;linux-x64;osx-x64;osx-arm64</RuntimeIdentifiers>
<!-- A plain "dotnet publish" with no -r defaults $(RuntimeIdentifier) to
the host's own RID; see Directory.Build.props, which is where that
default actually has to live (PublishDir's RID suffix is decided
there, evaluated before this file, so a default set only here would
be too late for it). -->
<SelfContained>true</SelfContained>
<PublishSingleFile>true</PublishSingleFile>
<IncludeNativeLibrariesForSelfExtract>true</IncludeNativeLibrariesForSelfExtract>
<EnableCompressionInSingleFile>true</EnableCompressionInSingleFile>
<ImplicitUsings>enable</ImplicitUsings>
<AllowUnsafeBlocks>true</AllowUnsafeBlocks>
<Version>0.0.1</Version>
<ServerGarbageCollection>true</ServerGarbageCollection>
<!-- Server GC reserves one large heap range per logical processor. On
high-core-count hosts those ranges can overlap fixed PS5 image bases
before the guest address space is established. Workstation GC avoids
that collision and keeps the emulator's required mappings available. -->
<ServerGarbageCollection>false</ServerGarbageCollection>
<ConcurrentGarbageCollection>true</ConcurrentGarbageCollection>
<TieredPGO>true</TieredPGO>
</PropertyGroup>
<!-- Background GC's write-watch revisit calls FlushProcessWriteBuffers,
which on macOS uses thread_get_register_pointer_values; under Rosetta 2
that Mach call can stall indefinitely on threads executing translated
guest code, wedging the whole runtime (every allocating thread then
blocks behind the never-finishing GC). Non-concurrent GC never takes
that path. Windows and Linux keep concurrent GC. -->
<PropertyGroup Condition="$([System.String]::Copy('$(RuntimeIdentifier)').StartsWith('osx'))">
<ConcurrentGarbageCollection>false</ConcurrentGarbageCollection>
</PropertyGroup>
<PropertyGroup Condition="'$(Configuration)' == 'Release'">
<GenerateDocumentationFile>false</GenerateDocumentationFile>
<DebugType>none</DebugType>
<DebugSymbols>false</DebugSymbols>
<RestorePackagesWithLockFile>true</RestorePackagesWithLockFile>
</PropertyGroup>
<PropertyGroup Condition="'$(RuntimeIdentifier)' == 'win-x64' Or '$(RuntimeIdentifier)' == ''">
<ApplicationIcon>..\..\assets\images\SharpEmu.ico</ApplicationIcon>
<Win32Icon>..\..\assets\images\SharpEmu.ico</Win32Icon>
<ApplicationManifest>app.manifest</ApplicationManifest>
</PropertyGroup>
<PropertyGroup>
<NoWarn>$(NoWarn);1591</NoWarn>
</PropertyGroup>
<ItemGroup>
<Content Include="..\..\LICENSE.txt">
<CopyToOutputDirectory>Always</CopyToOutputDirectory>
@@ -58,17 +79,75 @@ SPDX-License-Identifier: GPL-2.0-or-later
</Content>
</ItemGroup>
<!-- Keep glfw as a loose file next to the executable; every other native
<!-- Native libraries (glfw, FFmpeg) publish into a subfolder next to the
executable instead of sitting loose beside it, so the publish
directory stays uncluttered as more native deps get added. The folder
name is a fixed constant, not derived from the RID/architecture: each
publish output only ever holds one architecture's binaries anyway, so
varying the name added a class of bugs (RID resolution timing, host-OS
vs. target-RID mixups) for no benefit. Runtime code (Program.cs's
PreloadGlfw, FfmpegNativeBinkFrameSource's RootPath) uses the same
literal "plugins" folder name. -->
<PropertyGroup>
<NativeLibraryFolderName>plugins</NativeLibraryFolderName>
</PropertyGroup>
<!-- Keep glfw as a loose file in the native subfolder; every other native
library is embedded into the single-file bundle. -->
<Target Name="KeepGlfwOutsideSingleFile" AfterTargets="ComputeResolvedFilesToPublishList">
<ItemGroup>
<_GlfwPublishFiles Include="@(ResolvedFileToPublish)"
Condition="$([System.String]::Copy('%(ResolvedFileToPublish.Filename)').StartsWith('glfw'))" />
Condition="$([System.String]::Copy('%(ResolvedFileToPublish.Filename)').StartsWith('glfw')) Or $([System.String]::Copy('%(ResolvedFileToPublish.Filename)').StartsWith('libglfw'))" />
<ResolvedFileToPublish Remove="@(_GlfwPublishFiles)" />
<ResolvedFileToPublish Include="@(_GlfwPublishFiles)">
<ExcludeFromSingleFile>true</ExcludeFromSingleFile>
<RelativePath>$(NativeLibraryFolderName)/%(Filename)%(Extension)</RelativePath>
</ResolvedFileToPublish>
</ItemGroup>
</Target>
<PropertyGroup>
<FfmpegRuntimeTag>2c92585</FfmpegRuntimeTag>
<FfmpegRuntimeDir>
$(BaseIntermediateOutputPath)ffmpeg-runtime/$(FfmpegRuntimeTag)/$(RuntimeIdentifier)</FfmpegRuntimeDir>
<FfmpegRuntimePackage Condition="'$(RuntimeIdentifier)' == 'win-x64'">ffmpeg-windows-x64.zip</FfmpegRuntimePackage>
<FfmpegRuntimePackage Condition="'$(RuntimeIdentifier)' == 'linux-x64'">ffmpeg-linux-x64.zip</FfmpegRuntimePackage>
<FfmpegRuntimePackage Condition="'$(RuntimeIdentifier)' == 'osx-x64'">ffmpeg-macos-x64.zip</FfmpegRuntimePackage>
<FfmpegRuntimePackage Condition="'$(RuntimeIdentifier)' == 'osx-arm64'">ffmpeg-macos-arm64.zip</FfmpegRuntimePackage>
<FfmpegRuntimeArchive>$(FfmpegRuntimeDir)/$(FfmpegRuntimePackage)</FfmpegRuntimeArchive>
<FfmpegRuntimeExtractDir>$(FfmpegRuntimeDir)/extracted</FfmpegRuntimeExtractDir>
</PropertyGroup>
<Target Name="FetchFfmpegRuntime"
BeforeTargets="Publish"
Condition="'$(RuntimeIdentifier)' != '' And '$(FfmpegRuntimePackage)' != ''">
<DownloadFile
SourceUrl="https://github.com/sharpemu/ffmpeg-core/releases/download/$(FfmpegRuntimeTag)/$(FfmpegRuntimePackage)"
DestinationFolder="$(FfmpegRuntimeDir)"
Condition="!Exists('$(FfmpegRuntimeArchive)')" />
<Unzip
SourceFiles="$(FfmpegRuntimeArchive)"
DestinationFolder="$(FfmpegRuntimeExtractDir)"
Condition="!Exists('$(FfmpegRuntimeExtractDir)')" />
</Target>
<Target Name="PublishFfmpegRuntime"
AfterTargets="Publish"
DependsOnTargets="FetchFfmpegRuntime"
Condition="'$(RuntimeIdentifier)' != '' And '$(FfmpegRuntimePackage)' != ''">
<!-- Keyed off the target $(RuntimeIdentifier), not the host OS: publishing
e.g. linux-x64 from a Windows machine is a supported cross-publish,
and the extracted archive's own layout (bin/*.dll vs lib/*.so*) only
depends on which platform's ffmpeg-core package was fetched. -->
<ItemGroup>
<_FfmpegRuntimeFiles Condition="$(RuntimeIdentifier.StartsWith('win'))"
Include="$(FfmpegRuntimeExtractDir)/bin/*.dll" />
<_FfmpegRuntimeFiles Condition="!$(RuntimeIdentifier.StartsWith('win'))"
Include="$(FfmpegRuntimeExtractDir)/lib/*.so;$(FfmpegRuntimeExtractDir)/lib/*.so.*;$(FfmpegRuntimeExtractDir)/lib/*.dylib" />
</ItemGroup>
<Copy SourceFiles="@(_FfmpegRuntimeFiles)"
DestinationFolder="$(PublishDir)$(NativeLibraryFolderName)"
SkipUnchangedFiles="true" />
</Target>
</Project>
+21
View File
@@ -0,0 +1,21 @@
<?xml version="1.0" encoding="utf-8"?>
<!--
Copyright (C) 2026 SharpEmu Emulator Project
SPDX-License-Identifier: GPL-2.0-or-later
-->
<assembly manifestVersion="1.0" xmlns="urn:schemas-microsoft-com:asm.v1">
<assemblyIdentity version="1.0.0.0" name="SharpEmu" />
<compatibility xmlns="urn:schemas-microsoft-com:compatibility.v1">
<application>
<!-- Required by Avalonia NativeControlHost on Windows 10 and 11. -->
<supportedOS Id="{8e0f7a12-bfb3-4fe8-b9a5-48fd50a15a9a}" />
</application>
</compatibility>
<asmv3:application xmlns:asmv3="urn:schemas-microsoft-com:asm.v3">
<asmv3:windowsSettings>
<dpiAwareness xmlns="http://schemas.microsoft.com/SMI/2016/WindowsSettings">PerMonitorV2</dpiAwareness>
</asmv3:windowsSettings>
</asmv3:application>
</assembly>
-491
View File
@@ -1,491 +0,0 @@
{
"version": 2,
"dependencies": {
"net10.0": {
"Microsoft.NET.ILLink.Tasks": {
"type": "Direct",
"requested": "[10.0.3, )",
"resolved": "10.0.3",
"contentHash": "0B6nZyCHWXnvmlB559oduOspVdNOnpNXPjhpWVMovLPAsDVG7A4jJR9rzECf67JUzxP8/ee/wA8clwIzJcWNFA=="
},
"Avalonia.Angle.Windows.Natives": {
"type": "Transitive",
"resolved": "2.1.25547.20250602",
"contentHash": "ZL0VLc4s9rvNNFt19Pxm5UNAkmKNylugAwJPX9ulXZ6JWs/l6XZihPWWTyezaoNOVyEPU8YbURtW7XMAtqXH5A=="
},
"Avalonia.BuildServices": {
"type": "Transitive",
"resolved": "11.3.2",
"contentHash": "qHDToxto1e3hci5YqbG9n0Ty8mlp3zBUN5wT66wKqaDVzXyQ0do3EnRILd4Ke9jpvsktaPpgE0YjEk7hornryQ=="
},
"Avalonia.FreeDesktop": {
"type": "Transitive",
"resolved": "11.3.18",
"contentHash": "aUwv8BNruRUOaUfMu4U3uibIUS60/rSHgGOhd8zBkLkpxY3JFJvgRbeq5ZzHIyKXCuKi18PO00YHAgCarp3wdw==",
"dependencies": {
"Avalonia": "11.3.18",
"Tmds.DBus.Protocol": "0.21.3"
}
},
"Avalonia.Native": {
"type": "Transitive",
"resolved": "11.3.18",
"contentHash": "8g53DROFW6wVJAnTsE1Iu5bdZO5r0oqxbdbMQODs1QDYCK2IU/y/x/vQVKCYOvqxl4tXkzU9tExv8XqSGPWthQ==",
"dependencies": {
"Avalonia": "11.3.18"
}
},
"Avalonia.Remote.Protocol": {
"type": "Transitive",
"resolved": "11.3.18",
"contentHash": "vw+6ZfgTuu72dA9aVWn6u56t2nrBd5MoMU0wo/qI9XJAl/c0oYYphIvwLvJP1JorubQY4UE3d0ac8ULBhrGBiA=="
},
"Avalonia.Skia": {
"type": "Transitive",
"resolved": "11.3.18",
"contentHash": "/B4aXmNRNjG8I5U/a1xJI+bIi0XO6DDzS3mBrIKlVnJRY2CyZiUeESRQXLnIU77Z9TvqkUROs+D47s085YjFtA==",
"dependencies": {
"Avalonia": "11.3.18",
"HarfBuzzSharp": "8.3.1.1",
"HarfBuzzSharp.NativeAssets.Linux": "8.3.1.1",
"HarfBuzzSharp.NativeAssets.WebAssembly": "8.3.1.1",
"SkiaSharp": "2.88.9",
"SkiaSharp.NativeAssets.Linux": "2.88.9",
"SkiaSharp.NativeAssets.WebAssembly": "2.88.9"
}
},
"Avalonia.Win32": {
"type": "Transitive",
"resolved": "11.3.18",
"contentHash": "eioUHkM2PeLPETd1aEks3rvb9plbba6buIrNdrqCpwE/qgHKUjvRNBd5mUQfAbGgTLiAes524gB8uUMDhrsJVQ==",
"dependencies": {
"Avalonia": "11.3.18",
"Avalonia.Angle.Windows.Natives": "2.1.25547.20250602"
}
},
"Avalonia.X11": {
"type": "Transitive",
"resolved": "11.3.18",
"contentHash": "m4Ki/G5Dovnq+6QzfS0iGbK8V77Q6oTjToMLOB0CxPCCrl3Oxywh6kIjuGJDPaN6kopMmjxlNShyQf+vPYL+JA==",
"dependencies": {
"Avalonia": "11.3.18",
"Avalonia.FreeDesktop": "11.3.18",
"Avalonia.Skia": "11.3.18"
}
},
"HarfBuzzSharp": {
"type": "Transitive",
"resolved": "8.3.1.1",
"contentHash": "tLZN66oe/uiRPTZfrCU4i8ScVGwqHNh5MHrXj0yVf4l7Mz0FhTGnQ71RGySROTmdognAs0JtluHkL41pIabWuQ==",
"dependencies": {
"HarfBuzzSharp.NativeAssets.Win32": "8.3.1.1",
"HarfBuzzSharp.NativeAssets.macOS": "8.3.1.1"
}
},
"HarfBuzzSharp.NativeAssets.Linux": {
"type": "Transitive",
"resolved": "8.3.1.1",
"contentHash": "3EZ1mpIiKWRLL5hUYA82ZHteeDIVaEA/Z0rA/wU6tjx6crcAkJnBPwDXZugBSfo8+J3EznvRJf49uMsqYfKrHg=="
},
"HarfBuzzSharp.NativeAssets.macOS": {
"type": "Transitive",
"resolved": "8.3.1.1",
"contentHash": "jbtCsgftcaFLCA13tVKo5iWdElJScrulLTKJre36O4YQTIlwDtPPqhRZNk+Y0vv4D1gxbscasGRucUDfS44ofQ=="
},
"HarfBuzzSharp.NativeAssets.WebAssembly": {
"type": "Transitive",
"resolved": "8.3.1.1",
"contentHash": "loJweK2u/mH/3C2zBa0ggJlITIszOkK64HLAZB7FUT670dTg965whLFYHDQo69NmC4+d9UN0icLC9VHidXaVCA=="
},
"HarfBuzzSharp.NativeAssets.Win32": {
"type": "Transitive",
"resolved": "8.3.1.1",
"contentHash": "UsJtQsfAJoFDZrXc4hCUfRPMqccfKZ0iumJ/upcUjz/cmsTgVFGNEL5yaJWmkqsuFYdMWbj/En5/kS4PFl9hBA=="
},
"MicroCom.Runtime": {
"type": "Transitive",
"resolved": "0.11.0",
"contentHash": "MEnrZ3UIiH40hjzMDsxrTyi8dtqB5ziv3iBeeU4bXsL/7NLSal9F1lZKpK+tfBRnUoDSdtcW3KufE4yhATOMCA=="
},
"Microsoft.DotNet.PlatformAbstractions": {
"type": "Transitive",
"resolved": "3.1.6",
"contentHash": "jek4XYaQ/PGUwDKKhwR8K47Uh1189PFzMeLqO83mXrXQVIpARZCcfuDedH50YDTepBkfijCZN5U/vZi++erxtg=="
},
"Microsoft.Extensions.DependencyModel": {
"type": "Transitive",
"resolved": "9.0.9",
"contentHash": "fNGvKct2De8ghm0Bpfq0iWthtzIWabgOTi+gJhNOPhNJIowXNEUE2eZNW/zNCzrHVA3PXg2yZ+3cWZndC2IqYA=="
},
"Silk.NET.Core": {
"type": "Transitive",
"resolved": "2.23.0",
"contentHash": "D7AT/nnwlB+4RZ84XY8QNGBZMJI5z9l4CSSETIJ1wCfRJzRt/341y3MRZ4HbnFz4r/IGaWOEZr86iE+0/65yyQ==",
"dependencies": {
"Microsoft.DotNet.PlatformAbstractions": "3.1.6",
"Microsoft.Extensions.DependencyModel": "9.0.9"
}
},
"Silk.NET.GLFW": {
"type": "Transitive",
"resolved": "2.23.0",
"contentHash": "UIs4sH57xlPUNHQ/1bt9rymPWlGy8IMDCNv86h0iM4TOA1CkIx0XM/n/tA4AReh1zQkNrvkxPEdZ3Blvy1dyXg==",
"dependencies": {
"Silk.NET.Core": "2.23.0",
"Ultz.Native.GLFW": "3.4.0"
}
},
"Silk.NET.Maths": {
"type": "Transitive",
"resolved": "2.23.0",
"contentHash": "r8PdIVzME8EH0qAgbmRPO87I4GfgR2j8TofT7EMuRJDf1QluoQwnVypDoFJjQ2ZBSRsGYk5unYxxogI05Ogsmw=="
},
"Silk.NET.Windowing.Common": {
"type": "Transitive",
"resolved": "2.23.0",
"contentHash": "ThStSinmY9KQI8DGiF5XEhkLJVnBcgRTBTzL9ijg1wMZAYuckz7ykrNw04fjRm2Gryh6tCNGbvz2XaY0efeFzg==",
"dependencies": {
"Silk.NET.Core": "2.23.0",
"Silk.NET.Maths": "2.23.0"
}
},
"Silk.NET.Windowing.Glfw": {
"type": "Transitive",
"resolved": "2.23.0",
"contentHash": "aYBudKmENmvLRn9p15HbdvlQTnnXskcDfTfbYwSb/4fr263rGLwYuDw/txUEc2jihHJiWCp5+75Y7z5wTJWl7g==",
"dependencies": {
"Silk.NET.GLFW": "2.23.0",
"Silk.NET.Windowing.Common": "2.23.0"
}
},
"SkiaSharp": {
"type": "Transitive",
"resolved": "2.88.9",
"contentHash": "3MD5VHjXXieSHCleRLuaTXmL2pD0mB7CcOB1x2kA1I4bhptf4e3R27iM93264ZYuAq6mkUyX5XbcxnZvMJYc1Q==",
"dependencies": {
"SkiaSharp.NativeAssets.Win32": "2.88.9",
"SkiaSharp.NativeAssets.macOS": "2.88.9"
}
},
"SkiaSharp.NativeAssets.Linux": {
"type": "Transitive",
"resolved": "2.88.9",
"contentHash": "cWSaJKVPWAaT/WIn9c8T5uT/l4ETwHxNJTkEOtNKjphNo8AW6TF9O32aRkxqw3l8GUdUo66Bu7EiqtFh/XG0Zg==",
"dependencies": {
"SkiaSharp": "2.88.9"
}
},
"SkiaSharp.NativeAssets.macOS": {
"type": "Transitive",
"resolved": "2.88.9",
"contentHash": "Nv5spmKc4505Ep7oUoJ5vp3KweFpeNqxpyGDWyeEPTX2uR6S6syXIm3gj75dM0YJz7NPvcix48mR5laqs8dPuA=="
},
"SkiaSharp.NativeAssets.WebAssembly": {
"type": "Transitive",
"resolved": "2.88.9",
"contentHash": "kt06RccBHSnAs2wDYdBSfsjIDbY3EpsOVqnlDgKdgvyuRA8ZFDaHRdWNx1VHjGgYzmnFCGiTJBnXFl5BqGwGnA=="
},
"SkiaSharp.NativeAssets.Win32": {
"type": "Transitive",
"resolved": "2.88.9",
"contentHash": "wb2kYgU7iy84nQLYZwMeJXixvK++GoIuECjU4ECaUKNuflyRlJKyiRhN1MAHswvlvzuvkrjRWlK0Za6+kYQK7w=="
},
"Ultz.Native.GLFW": {
"type": "Transitive",
"resolved": "3.4.0",
"contentHash": "Iy22JopynbOJ32vA0lBhFEzGi65GQJBuJHYBYRBpydrDpNoTiHnjIXfA65Gu+8qsOr/ZEoIF8r9aHCgAXuO6DA=="
},
"sharpemu.core": {
"type": "Project",
"dependencies": {
"Iced": "[1.21.0, )",
"SharpEmu.HLE": "[1.0.0, )",
"SharpEmu.Libs": "[1.0.0, )",
"SharpEmu.Logging": "[1.0.0, )"
}
},
"sharpemu.gui": {
"type": "Project",
"dependencies": {
"Avalonia": "[11.3.18, )",
"Avalonia.Desktop": "[11.3.18, )",
"Avalonia.Fonts.Inter": "[11.3.18, )",
"Avalonia.Themes.Fluent": "[11.3.18, )",
"SharpEmu.Logging": "[1.0.0, )",
"Tmds.DBus.Protocol": "[0.21.3, )"
}
},
"sharpemu.hle": {
"type": "Project",
"dependencies": {
"SharpEmu.Logging": "[1.0.0, )"
}
},
"sharpemu.libs": {
"type": "Project",
"dependencies": {
"SharpEmu.HLE": "[1.0.0, )",
"Silk.NET.Vulkan": "[2.23.0, )",
"Silk.NET.Vulkan.Extensions.EXT": "[2.23.0, )",
"Silk.NET.Vulkan.Extensions.KHR": "[2.23.0, )",
"Silk.NET.Windowing": "[2.23.0, )"
}
},
"sharpemu.logging": {
"type": "Project"
},
"Avalonia": {
"type": "CentralTransitive",
"requested": "[11.3.18, )",
"resolved": "11.3.18",
"contentHash": "2C4UxhWUObWGgYKWic1x5BMMWGJP6SElb91WeOxs+X/iR26rtkqpxFFwwo50FXS9AyYnHfk8QKXDEfe7oT/kZA==",
"dependencies": {
"Avalonia.BuildServices": "11.3.2",
"Avalonia.Remote.Protocol": "11.3.18",
"MicroCom.Runtime": "0.11.0"
}
},
"Avalonia.Desktop": {
"type": "CentralTransitive",
"requested": "[11.3.18, )",
"resolved": "11.3.18",
"contentHash": "bilMPa5vYiis6fbNovb6esKytBnOCEGojBa1XFegLCRHCP6g6PvZwS0XF/YOAGkENRlHG8dI7lohOpQ9bIkq1g==",
"dependencies": {
"Avalonia": "11.3.18",
"Avalonia.Native": "11.3.18",
"Avalonia.Skia": "11.3.18",
"Avalonia.Win32": "11.3.18",
"Avalonia.X11": "11.3.18"
}
},
"Avalonia.Fonts.Inter": {
"type": "CentralTransitive",
"requested": "[11.3.18, )",
"resolved": "11.3.18",
"contentHash": "27u6hB3Y2Ue586yjfeVakberY73VNQXtuKwe/P927XG1QPlhsfmOyifLHDDpSHG85Zl1x/Xv9IZ3+tk9FnjcZQ==",
"dependencies": {
"Avalonia": "11.3.18"
}
},
"Avalonia.Themes.Fluent": {
"type": "CentralTransitive",
"requested": "[11.3.18, )",
"resolved": "11.3.18",
"contentHash": "+Q/TJoynD0zNuu5w2gD+xcTl7GNKJFxlPYAndRLs/mTDrNbbsvv/271WyIysbMPsXSjCyBDp7RCZzQkpD6x5Bg==",
"dependencies": {
"Avalonia": "11.3.18"
}
},
"Iced": {
"type": "CentralTransitive",
"requested": "[1.21.0, )",
"resolved": "1.21.0",
"contentHash": "dv5+81Q1TBQvVMSOOOmRcjJmvWcX3BZPZsIq31+RLc5cNft0IHAyNlkdb7ZarOWG913PyBoFDsDXoCIlKmLclg=="
},
"Silk.NET.Vulkan": {
"type": "CentralTransitive",
"requested": "[2.23.0, )",
"resolved": "2.23.0",
"contentHash": "3/irtlSWXZ3eTi8N6nelI6L34NTB8ZJHpqVMNzZx2aX7Ek9YEQ34NoQW8/Tljrtmkg8KRhHW8hKTEzZaKV8PgA==",
"dependencies": {
"Silk.NET.Core": "2.23.0"
}
},
"Silk.NET.Vulkan.Extensions.EXT": {
"type": "CentralTransitive",
"requested": "[2.23.0, )",
"resolved": "2.23.0",
"contentHash": "+Oth189ksRiL6HvGCwIdnsYHawqrbO8y49u1H61z3wsfcHhQZeVDYe/wF5LD7fk3NcdgDvwFD3mLm1QWhdZySw==",
"dependencies": {
"Silk.NET.Core": "2.23.0",
"Silk.NET.Vulkan": "2.23.0"
}
},
"Silk.NET.Vulkan.Extensions.KHR": {
"type": "CentralTransitive",
"requested": "[2.23.0, )",
"resolved": "2.23.0",
"contentHash": "uRaf4j+SmH3DumjSSSUbFg33BnsGZUyXGj93O9NgGKZSJN3OTmNmQDxRew+/KiVLcgH6qzbto8aNGZ++j9GFWg==",
"dependencies": {
"Silk.NET.Core": "2.23.0",
"Silk.NET.Vulkan": "2.23.0"
}
},
"Silk.NET.Windowing": {
"type": "CentralTransitive",
"requested": "[2.23.0, )",
"resolved": "2.23.0",
"contentHash": "OPNPmt/lRyUKVYrFLQXVxyATqD3MKLc1iY1oKx1/2GppgmZxVZPwN12tekrQ4C7408kgB1L5JD1Wnirqqeb2kg==",
"dependencies": {
"Silk.NET.Windowing.Common": "2.23.0",
"Silk.NET.Windowing.Glfw": "2.23.0"
}
},
"Tmds.DBus.Protocol": {
"type": "CentralTransitive",
"requested": "[0.21.3, )",
"resolved": "0.21.3",
"contentHash": "hDwB8WsQoyALQKqIbwzS68UKdlnafDm4T/DkO/JrA/YIneP/rKv96SxYPVXeh3FP4i/SXfShrYftKLtciJAIlw=="
}
},
"net10.0/linux-x64": {
"Avalonia.Angle.Windows.Natives": {
"type": "Transitive",
"resolved": "2.1.25547.20250602",
"contentHash": "ZL0VLc4s9rvNNFt19Pxm5UNAkmKNylugAwJPX9ulXZ6JWs/l6XZihPWWTyezaoNOVyEPU8YbURtW7XMAtqXH5A=="
},
"Avalonia.Native": {
"type": "Transitive",
"resolved": "11.3.18",
"contentHash": "8g53DROFW6wVJAnTsE1Iu5bdZO5r0oqxbdbMQODs1QDYCK2IU/y/x/vQVKCYOvqxl4tXkzU9tExv8XqSGPWthQ==",
"dependencies": {
"Avalonia": "11.3.18"
}
},
"HarfBuzzSharp.NativeAssets.Linux": {
"type": "Transitive",
"resolved": "8.3.1.1",
"contentHash": "3EZ1mpIiKWRLL5hUYA82ZHteeDIVaEA/Z0rA/wU6tjx6crcAkJnBPwDXZugBSfo8+J3EznvRJf49uMsqYfKrHg=="
},
"HarfBuzzSharp.NativeAssets.macOS": {
"type": "Transitive",
"resolved": "8.3.1.1",
"contentHash": "jbtCsgftcaFLCA13tVKo5iWdElJScrulLTKJre36O4YQTIlwDtPPqhRZNk+Y0vv4D1gxbscasGRucUDfS44ofQ=="
},
"HarfBuzzSharp.NativeAssets.Win32": {
"type": "Transitive",
"resolved": "8.3.1.1",
"contentHash": "UsJtQsfAJoFDZrXc4hCUfRPMqccfKZ0iumJ/upcUjz/cmsTgVFGNEL5yaJWmkqsuFYdMWbj/En5/kS4PFl9hBA=="
},
"SkiaSharp.NativeAssets.Linux": {
"type": "Transitive",
"resolved": "2.88.9",
"contentHash": "cWSaJKVPWAaT/WIn9c8T5uT/l4ETwHxNJTkEOtNKjphNo8AW6TF9O32aRkxqw3l8GUdUo66Bu7EiqtFh/XG0Zg==",
"dependencies": {
"SkiaSharp": "2.88.9"
}
},
"SkiaSharp.NativeAssets.macOS": {
"type": "Transitive",
"resolved": "2.88.9",
"contentHash": "Nv5spmKc4505Ep7oUoJ5vp3KweFpeNqxpyGDWyeEPTX2uR6S6syXIm3gj75dM0YJz7NPvcix48mR5laqs8dPuA=="
},
"SkiaSharp.NativeAssets.Win32": {
"type": "Transitive",
"resolved": "2.88.9",
"contentHash": "wb2kYgU7iy84nQLYZwMeJXixvK++GoIuECjU4ECaUKNuflyRlJKyiRhN1MAHswvlvzuvkrjRWlK0Za6+kYQK7w=="
},
"Ultz.Native.GLFW": {
"type": "Transitive",
"resolved": "3.4.0",
"contentHash": "Iy22JopynbOJ32vA0lBhFEzGi65GQJBuJHYBYRBpydrDpNoTiHnjIXfA65Gu+8qsOr/ZEoIF8r9aHCgAXuO6DA=="
}
},
"net10.0/osx-arm64": {
"Avalonia.Angle.Windows.Natives": {
"type": "Transitive",
"resolved": "2.1.25547.20250602",
"contentHash": "ZL0VLc4s9rvNNFt19Pxm5UNAkmKNylugAwJPX9ulXZ6JWs/l6XZihPWWTyezaoNOVyEPU8YbURtW7XMAtqXH5A=="
},
"Avalonia.Native": {
"type": "Transitive",
"resolved": "11.3.18",
"contentHash": "8g53DROFW6wVJAnTsE1Iu5bdZO5r0oqxbdbMQODs1QDYCK2IU/y/x/vQVKCYOvqxl4tXkzU9tExv8XqSGPWthQ==",
"dependencies": {
"Avalonia": "11.3.18"
}
},
"HarfBuzzSharp.NativeAssets.Linux": {
"type": "Transitive",
"resolved": "8.3.1.1",
"contentHash": "3EZ1mpIiKWRLL5hUYA82ZHteeDIVaEA/Z0rA/wU6tjx6crcAkJnBPwDXZugBSfo8+J3EznvRJf49uMsqYfKrHg=="
},
"HarfBuzzSharp.NativeAssets.macOS": {
"type": "Transitive",
"resolved": "8.3.1.1",
"contentHash": "jbtCsgftcaFLCA13tVKo5iWdElJScrulLTKJre36O4YQTIlwDtPPqhRZNk+Y0vv4D1gxbscasGRucUDfS44ofQ=="
},
"HarfBuzzSharp.NativeAssets.Win32": {
"type": "Transitive",
"resolved": "8.3.1.1",
"contentHash": "UsJtQsfAJoFDZrXc4hCUfRPMqccfKZ0iumJ/upcUjz/cmsTgVFGNEL5yaJWmkqsuFYdMWbj/En5/kS4PFl9hBA=="
},
"SkiaSharp.NativeAssets.Linux": {
"type": "Transitive",
"resolved": "2.88.9",
"contentHash": "cWSaJKVPWAaT/WIn9c8T5uT/l4ETwHxNJTkEOtNKjphNo8AW6TF9O32aRkxqw3l8GUdUo66Bu7EiqtFh/XG0Zg==",
"dependencies": {
"SkiaSharp": "2.88.9"
}
},
"SkiaSharp.NativeAssets.macOS": {
"type": "Transitive",
"resolved": "2.88.9",
"contentHash": "Nv5spmKc4505Ep7oUoJ5vp3KweFpeNqxpyGDWyeEPTX2uR6S6syXIm3gj75dM0YJz7NPvcix48mR5laqs8dPuA=="
},
"SkiaSharp.NativeAssets.Win32": {
"type": "Transitive",
"resolved": "2.88.9",
"contentHash": "wb2kYgU7iy84nQLYZwMeJXixvK++GoIuECjU4ECaUKNuflyRlJKyiRhN1MAHswvlvzuvkrjRWlK0Za6+kYQK7w=="
},
"Ultz.Native.GLFW": {
"type": "Transitive",
"resolved": "3.4.0",
"contentHash": "Iy22JopynbOJ32vA0lBhFEzGi65GQJBuJHYBYRBpydrDpNoTiHnjIXfA65Gu+8qsOr/ZEoIF8r9aHCgAXuO6DA=="
}
},
"net10.0/win-x64": {
"Avalonia.Angle.Windows.Natives": {
"type": "Transitive",
"resolved": "2.1.25547.20250602",
"contentHash": "ZL0VLc4s9rvNNFt19Pxm5UNAkmKNylugAwJPX9ulXZ6JWs/l6XZihPWWTyezaoNOVyEPU8YbURtW7XMAtqXH5A=="
},
"Avalonia.Native": {
"type": "Transitive",
"resolved": "11.3.18",
"contentHash": "8g53DROFW6wVJAnTsE1Iu5bdZO5r0oqxbdbMQODs1QDYCK2IU/y/x/vQVKCYOvqxl4tXkzU9tExv8XqSGPWthQ==",
"dependencies": {
"Avalonia": "11.3.18"
}
},
"HarfBuzzSharp.NativeAssets.Linux": {
"type": "Transitive",
"resolved": "8.3.1.1",
"contentHash": "3EZ1mpIiKWRLL5hUYA82ZHteeDIVaEA/Z0rA/wU6tjx6crcAkJnBPwDXZugBSfo8+J3EznvRJf49uMsqYfKrHg=="
},
"HarfBuzzSharp.NativeAssets.macOS": {
"type": "Transitive",
"resolved": "8.3.1.1",
"contentHash": "jbtCsgftcaFLCA13tVKo5iWdElJScrulLTKJre36O4YQTIlwDtPPqhRZNk+Y0vv4D1gxbscasGRucUDfS44ofQ=="
},
"HarfBuzzSharp.NativeAssets.Win32": {
"type": "Transitive",
"resolved": "8.3.1.1",
"contentHash": "UsJtQsfAJoFDZrXc4hCUfRPMqccfKZ0iumJ/upcUjz/cmsTgVFGNEL5yaJWmkqsuFYdMWbj/En5/kS4PFl9hBA=="
},
"SkiaSharp.NativeAssets.Linux": {
"type": "Transitive",
"resolved": "2.88.9",
"contentHash": "cWSaJKVPWAaT/WIn9c8T5uT/l4ETwHxNJTkEOtNKjphNo8AW6TF9O32aRkxqw3l8GUdUo66Bu7EiqtFh/XG0Zg==",
"dependencies": {
"SkiaSharp": "2.88.9"
}
},
"SkiaSharp.NativeAssets.macOS": {
"type": "Transitive",
"resolved": "2.88.9",
"contentHash": "Nv5spmKc4505Ep7oUoJ5vp3KweFpeNqxpyGDWyeEPTX2uR6S6syXIm3gj75dM0YJz7NPvcix48mR5laqs8dPuA=="
},
"SkiaSharp.NativeAssets.Win32": {
"type": "Transitive",
"resolved": "2.88.9",
"contentHash": "wb2kYgU7iy84nQLYZwMeJXixvK++GoIuECjU4ECaUKNuflyRlJKyiRhN1MAHswvlvzuvkrjRWlK0Za6+kYQK7w=="
},
"Ultz.Native.GLFW": {
"type": "Transitive",
"resolved": "3.4.0",
"contentHash": "Iy22JopynbOJ32vA0lBhFEzGi65GQJBuJHYBYRBpydrDpNoTiHnjIXfA65Gu+8qsOr/ZEoIF8r9aHCgAXuO6DA=="
}
}
}
}
+108 -40
View File
@@ -3,33 +3,38 @@
using System.Buffers.Binary;
using System.Text;
using SharpEmu.Core.Cpu.Debugging;
using SharpEmu.Core.Cpu.Native;
using SharpEmu.Core.Loader;
using SharpEmu.Core.Memory;
using SharpEmu.HLE;
using SharpEmu.Logging;
namespace SharpEmu.Core.Cpu;
public sealed class CpuDispatcher : ICpuDispatcher, IDisposable
{
private static readonly SharpEmuLogger Log = SharpEmuLog.For("Dispatcher");
private enum EntryFrameKind
{
ProcessEntry,
ModuleInitializer,
}
private const ulong StackBaseAddress = 0x7FFF_F000_0000UL;
// The top of the x86-64 user address space (0x7FFD..0x7FFF) is only
// freely mappable on Windows; on macOS/Linux it hosts the dyld shared
// cache / vdso and (under Rosetta 2) the translator runtime, so POSIX
// hosts use the equivalent layout one slot lower at 0x6FFx.
private static readonly ulong StackBaseAddress = OperatingSystem.IsWindows() ? 0x7FFF_F000_0000UL : 0x6FFF_F000_0000UL;
private const ulong StackSize = 0x0020_0000UL;
private const ulong TlsBaseAddress = 0x7FFE_0000_0000UL;
private static readonly ulong TlsBaseAddress = OperatingSystem.IsWindows() ? 0x7FFE_0000_0000UL : 0x6FFE_0000_0000UL;
private const ulong TlsSize = 0x0001_0000UL;
private const ulong TlsPrefixSize = 0x0000_1000UL;
private const ulong BootstrapStubBaseAddress = 0x7FFD_F000_0000UL;
private const ulong BootstrapPayloadBaseAddress = 0x7FFD_E000_0000UL;
private const ulong DynlibFallbackStubBaseAddress = 0x7FFD_D000_0000UL;
private const ulong ReturnToHostStubBaseAddress = 0x7FFD_C000_0000UL;
// The static TLS blocks live at negative offsets from the TCB (FreeBSD
// amd64 variant II). Keep every host in sync with GuestTlsTemplate's
// startup reservation; PS5 modules routinely reach beyond one host page.
private const ulong TlsPrefixSize = GuestTlsTemplate.StartupStaticTlsReservation;
private static readonly ulong BootstrapStubBaseAddress = OperatingSystem.IsWindows() ? 0x7FFD_F000_0000UL : 0x6FFD_F000_0000UL;
private static readonly ulong BootstrapPayloadBaseAddress = OperatingSystem.IsWindows() ? 0x7FFD_E000_0000UL : 0x6FFD_E000_0000UL;
private static readonly ulong DynlibFallbackStubBaseAddress = OperatingSystem.IsWindows() ? 0x7FFD_D000_0000UL : 0x6FFD_D000_0000UL;
private static readonly ulong ReturnToHostStubBaseAddress = OperatingSystem.IsWindows() ? 0x7FFD_C000_0000UL : 0x6FFD_C000_0000UL;
private const ulong BootstrapRegionSize = 0x0000_1000UL;
private const ulong ReturnToHostStubStride = 0x0100_0000UL;
private const ulong BootstrapPayloadResultOffset = 0x28UL;
@@ -83,8 +88,8 @@ public sealed class CpuDispatcher : ICpuDispatcher, IDisposable
string processImageName = "eboot.bin",
CpuExecutionOptions executionOptions = default)
{
Log.Debug("=== DispatchEntry START ===");
Log.Debug($"entryPoint=0x{entryPoint:X16}, generation={generation}");
Console.Error.WriteLine("[DISPATCHER] === DispatchEntry START ===");
Console.Error.WriteLine($"[DISPATCHER] entryPoint=0x{entryPoint:X16}, generation={generation}");
try
{
@@ -92,7 +97,8 @@ public sealed class CpuDispatcher : ICpuDispatcher, IDisposable
}
catch (Exception ex)
{
Log.Critical($"FATAL EXCEPTION in DispatchEntry: {ex.GetType().Name}: {ex.Message}", ex);
Console.Error.WriteLine($"[DISPATCHER] FATAL EXCEPTION in DispatchEntry: {ex.GetType().Name}: {ex.Message}");
Console.Error.WriteLine($"[DISPATCHER] Stack trace: {ex.StackTrace}");
throw;
}
}
@@ -105,8 +111,8 @@ public sealed class CpuDispatcher : ICpuDispatcher, IDisposable
string moduleName = "module",
CpuExecutionOptions executionOptions = default)
{
Log.Debug("=== DispatchModuleInitializer START ===");
Log.Debug($"moduleInit=0x{entryPoint:X16}, generation={generation}, module={moduleName}");
Console.Error.WriteLine("[DISPATCHER] === DispatchModuleInitializer START ===");
Console.Error.WriteLine($"[DISPATCHER] moduleInit=0x{entryPoint:X16}, generation={generation}, module={moduleName}");
try
{
@@ -121,7 +127,8 @@ public sealed class CpuDispatcher : ICpuDispatcher, IDisposable
}
catch (Exception ex)
{
Log.Critical($"FATAL EXCEPTION in DispatchModuleInitializer: {ex.GetType().Name}: {ex.Message}", ex);
Console.Error.WriteLine($"[DISPATCHER] FATAL EXCEPTION in DispatchModuleInitializer: {ex.GetType().Name}: {ex.Message}");
Console.Error.WriteLine($"[DISPATCHER] Stack trace: {ex.StackTrace}");
throw;
}
}
@@ -135,7 +142,7 @@ public sealed class CpuDispatcher : ICpuDispatcher, IDisposable
CpuExecutionOptions executionOptions = default,
EntryFrameKind frameKind = EntryFrameKind.ProcessEntry)
{
Log.Debug("DispatchEntryCore STARTING...");
Console.Error.WriteLine("[DISPATCHER] DispatchEntryCore STARTING...");
LastEntryPoint = entryPoint;
LastTrapInfo = null;
@@ -266,7 +273,23 @@ public sealed class CpuDispatcher : ICpuDispatcher, IDisposable
entryFrameDiagnostic,
Environment.NewLine,
"CpuEngine: native-only");
// Frame boundaries an attached debugger observes; null hook = a branch.
var debugHook = executionOptions.DebugHook;
var debugFrame = debugHook is null
? null
: new CpuContextDebugFrame(
frameKind == EntryFrameKind.ProcessEntry
? CpuDebugFrameKind.ProcessEntry
: CpuDebugFrameKind.ModuleInitializer,
entryPoint,
processImageName,
context,
effectiveImportStubs);
debugHook?.OnFrameEnter(debugFrame!);
_nativeCpuBackend ??= new DirectExecutionBackend(_moduleManager);
// Let backend stall reports reference the same frame as entry.
(_nativeCpuBackend as DirectExecutionBackend)?.SetActiveDebugFrame(debugFrame);
if (_nativeCpuBackend.TryExecute(
context,
entryPoint,
@@ -276,6 +299,7 @@ public sealed class CpuDispatcher : ICpuDispatcher, IDisposable
executionOptions,
out var nativeResult))
{
debugHook?.OnFrameExit(debugFrame!, nativeResult);
LastSessionSummary = new CpuSessionSummary(
nativeResult,
nativeResult == OrbisGen2Result.ORBIS_GEN2_OK
@@ -290,6 +314,8 @@ public sealed class CpuDispatcher : ICpuDispatcher, IDisposable
return nativeResult;
}
debugHook?.OnFrameExit(debugFrame!, OrbisGen2Result.ORBIS_GEN2_ERROR_NOT_IMPLEMENTED);
var backendName = string.IsNullOrWhiteSpace(_nativeCpuBackend.BackendName)
? "native-backend"
: _nativeCpuBackend.BackendName;
@@ -307,7 +333,7 @@ public sealed class CpuDispatcher : ICpuDispatcher, IDisposable
LastMilestoneLog,
Environment.NewLine,
$"CpuEngine native-only failed: {backendError}");
Log.Error($"Native backend FAILED: {backendError}");
Console.Error.WriteLine($"[DISPATCHER] Native backend FAILED: {backendError}");
return FailEarly(
OrbisGen2Result.ORBIS_GEN2_ERROR_NOT_IMPLEMENTED,
CpuExitReason.NativeBackendUnavailable);
@@ -366,11 +392,19 @@ public sealed class CpuDispatcher : ICpuDispatcher, IDisposable
private static bool InitializeTls(CpuContext context, ulong tlsBase)
{
return context.TryWriteUInt64(tlsBase - 0xF0, 0) &&
context.TryWriteUInt64(tlsBase + 0x00, tlsBase) &&
context.TryWriteUInt64(tlsBase + 0x10, tlsBase) &&
context.TryWriteUInt64(tlsBase + 0x28, 0xC0DEC0DECAFEBABEUL) &&
context.TryWriteUInt64(tlsBase + 0x60, tlsBase);
if (!context.TryWriteUInt64(tlsBase - 0xF0, 0) ||
!context.TryWriteUInt64(tlsBase + 0x00, tlsBase) ||
!context.TryWriteUInt64(tlsBase + 0x10, tlsBase) ||
!context.TryWriteUInt64(tlsBase + 0x28, 0xC0DEC0DECAFEBA00UL) ||
!context.TryWriteUInt64(tlsBase + 0x60, tlsBase))
{
return false;
}
// Seed the static TLS block below the thread pointer with the main
// module's initialized thread-locals (variant II layout).
SharpEmu.HLE.GuestTlsTemplate.SeedThreadBlock(context, tlsBase);
return true;
}
private static bool InitializeGuestFrameChainSentinel(CpuContext context)
@@ -396,35 +430,54 @@ public sealed class CpuDispatcher : ICpuDispatcher, IDisposable
ulong programExitHandlerAddress)
{
var imageName = string.IsNullOrWhiteSpace(processImageName) ? "eboot.bin" : processImageName;
var encodedNameLength = Encoding.UTF8.GetByteCount(imageName);
Span<byte> argv0Buffer = encodedNameLength + 1 <= 512
? stackalloc byte[encodedNameLength + 1]
: new byte[encodedNameLength + 1];
if (Encoding.UTF8.GetBytes(imageName.AsSpan(), argv0Buffer) != encodedNameLength)
var arguments = new List<string>(3) { imageName };
var configuredArguments = Environment.GetEnvironmentVariable("SHARPEMU_GUEST_ARGS");
if (!string.IsNullOrWhiteSpace(configuredArguments))
{
return false;
// The PS5 entry-parameter ABI exposes three inline argv pointers.
// Two compatibility arguments are therefore safe without changing
// the fixed 0x20-byte structure expected by existing titles.
var compatibilityArguments = configuredArguments.Split(
(char[]?)null,
StringSplitOptions.RemoveEmptyEntries | StringSplitOptions.TrimEntries);
arguments.AddRange(compatibilityArguments.Take(2));
}
argv0Buffer[encodedNameLength] = 0;
var cursor = context[CpuRegister.Rsp];
var argv0Address = AlignDown(cursor - (ulong)argv0Buffer.Length, 16);
if (!context.Memory.TryWrite(argv0Address, argv0Buffer))
var argumentAddresses = new ulong[arguments.Count];
for (var index = arguments.Count - 1; index >= 0; index--)
{
return false;
var encoded = Encoding.UTF8.GetBytes(arguments[index] + '\0');
cursor = AlignDown(cursor - (ulong)encoded.Length, 16);
if (!context.Memory.TryWrite(cursor, encoded))
{
return false;
}
argumentAddresses[index] = cursor;
}
const ulong entryParamsSize = 0x20;
var entryParamsAddress = AlignDown(argv0Address - entryParamsSize, 16);
if (!context.TryWriteUInt32(entryParamsAddress + 0x00, 1) ||
!context.TryWriteUInt32(entryParamsAddress + 0x04, 0) ||
!context.TryWriteUInt64(entryParamsAddress + 0x08, argv0Address) ||
!context.TryWriteUInt64(entryParamsAddress + 0x10, 0) ||
!context.TryWriteUInt64(entryParamsAddress + 0x18, 0))
var entryParamsAddress = AlignDown(cursor - entryParamsSize, 16);
if (!TryWriteUInt32(context, entryParamsAddress + 0x00, (uint)arguments.Count) ||
!TryWriteUInt32(context, entryParamsAddress + 0x04, 0) ||
!context.TryWriteUInt64(entryParamsAddress + 0x08, argumentAddresses[0]) ||
!context.TryWriteUInt64(
entryParamsAddress + 0x10,
argumentAddresses.Length > 1 ? argumentAddresses[1] : 0) ||
!context.TryWriteUInt64(
entryParamsAddress + 0x18,
argumentAddresses.Length > 2 ? argumentAddresses[2] : 0))
{
return false;
}
if (arguments.Count > 1)
{
Console.Error.WriteLine(
$"[DISPATCHER] Guest arguments: {string.Join(' ', arguments.Skip(1))}");
}
var entryStackPointer = entryParamsAddress - sizeof(ulong);
if (!context.TryWriteUInt64(entryStackPointer, 0))
{
@@ -457,6 +510,13 @@ public sealed class CpuDispatcher : ICpuDispatcher, IDisposable
return value & ~(alignment - 1);
}
private static bool TryWriteUInt32(CpuContext context, ulong address, uint value)
{
Span<byte> buffer = stackalloc byte[sizeof(uint)];
BinaryPrimitives.WriteUInt32LittleEndian(buffer, value);
return context.Memory.TryWrite(address, buffer);
}
private static string BuildEntryFrameDiagnostic(
ulong entryPoint,
CpuContext context,
@@ -653,6 +713,13 @@ public sealed class CpuDispatcher : ICpuDispatcher, IDisposable
return _virtualMemory.TryWrite(address, buffer);
}
/// <summary>
/// True when the disposed native backend left its session state alive
/// because guest workers were still executing guest code. The guest
/// address space must then stay mapped as well.
/// </summary>
internal bool NativeSessionLeaked { get; private set; }
public void Dispose()
{
if (_nativeCpuBackend is IDisposable disposableBackend)
@@ -660,6 +727,7 @@ public sealed class CpuDispatcher : ICpuDispatcher, IDisposable
disposableBackend.Dispose();
}
NativeSessionLeaked = _nativeCpuBackend is DirectExecutionBackend { GuestSessionLeaked: true };
_nativeCpuBackend = null;
}
}
+12 -1
View File
@@ -1,15 +1,26 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using SharpEmu.Core.Cpu.Debugging;
namespace SharpEmu.Core.Cpu;
public readonly struct CpuExecutionOptions
{
public bool EnableDisasmDiagnostics { get; init; }
public CpuExecutionEngine CpuEngine { get; init; }
public bool StrictDynlibResolution { get; init; }
public int ImportTraceLimit { get; init; }
/// <summary>
/// An optional debugger attached to this execution session. When set, the
/// dispatcher notifies it at each frame boundary via
/// <see cref="ICpuDebugHook.OnFrameEnter"/> / <see cref="ICpuDebugHook.OnFrameExit"/>.
/// Null when no debugger is attached, which is the default and imposes no
/// runtime cost.
/// </summary>
public ICpuDebugHook? DebugHook { get; init; }
}
@@ -0,0 +1,66 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using SharpEmu.HLE;
namespace SharpEmu.Core.Cpu.Debugging;
/// <summary>
/// Adapts a live <see cref="CpuContext"/> to <see cref="ICpuDebugFrame"/>. The
/// dispatcher creates one of these around the guest context it is about to run
/// and passes it to the attached <see cref="ICpuDebugHook"/>; every accessor
/// forwards directly to the underlying context.
/// </summary>
internal sealed class CpuContextDebugFrame : ICpuDebugFrame
{
private readonly CpuContext _context;
internal CpuContextDebugFrame(
CpuDebugFrameKind kind,
ulong entryPoint,
string label,
CpuContext context,
IReadOnlyDictionary<ulong, string> importStubs)
{
Kind = kind;
EntryPoint = entryPoint;
Label = label ?? string.Empty;
_context = context ?? throw new ArgumentNullException(nameof(context));
ImportStubs = importStubs ?? new Dictionary<ulong, string>();
}
public CpuDebugFrameKind Kind { get; }
public Generation Generation => _context.TargetGeneration;
public ulong EntryPoint { get; }
public string Label { get; }
public ICpuMemory Memory => _context.Memory;
public ulong GetRegister(CpuRegister register) => _context[register];
public void SetRegister(CpuRegister register, ulong value) => _context[register] = value;
public ulong Rip
{
get => _context.Rip;
set => _context.Rip = value;
}
public ulong Rflags
{
get => _context.Rflags;
set => _context.Rflags = value;
}
public ulong FsBase => _context.FsBase;
public ulong GsBase => _context.GsBase;
public void GetXmm(int registerIndex, out ulong low, out ulong high)
=> _context.GetXmmRegister(registerIndex, out low, out high);
public IReadOnlyDictionary<ulong, string> ImportStubs { get; }
}
@@ -0,0 +1,18 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
namespace SharpEmu.Core.Cpu.Debugging;
/// <summary>
/// Identifies the kind of guest entry frame a debugger is observing. The
/// dispatcher enters a fresh frame for the process entry point and for every
/// module initializer, so the debug layer can label stops accordingly.
/// </summary>
public enum CpuDebugFrameKind
{
/// <summary>The guest process entry point (<c>eboot.bin</c> start).</summary>
ProcessEntry,
/// <summary>A module DT_INIT / initializer routine.</summary>
ModuleInitializer,
}
@@ -0,0 +1,70 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
namespace SharpEmu.Core.Cpu.Debugging;
/// <summary>The kind of execution stall the backend detected.</summary>
public enum CpuStallKind
{
/// <summary>
/// The guest is repeatedly re-dispatching the same import with no forward
/// progress — most commonly a spin on a mutex lock/unlock pair.
/// </summary>
ImportLoop,
}
/// <summary>
/// Details of a detected stall handed to <see cref="ICpuDebugHook.OnStall"/>.
/// Reported from the emulation thread at the point the backend recognises the
/// livelock, before it forces the guest out of the loop.
/// </summary>
public readonly struct CpuStallInfo
{
public CpuStallInfo(
CpuStallKind kind,
string? nid,
ulong instructionPointer,
long dispatchIndex,
ulong argument0,
ulong argument1,
string detail,
string? libraryName = null,
string? functionName = null)
{
Kind = kind;
Nid = nid;
InstructionPointer = instructionPointer;
DispatchIndex = dispatchIndex;
Argument0 = argument0;
Argument1 = argument1;
Detail = detail ?? string.Empty;
LibraryName = libraryName;
FunctionName = functionName;
}
public CpuStallKind Kind { get; }
/// <summary>The NID of the import being spun on, when known.</summary>
public string? Nid { get; }
/// <summary>The guest return address of the looping import dispatch.</summary>
public ulong InstructionPointer { get; }
/// <summary>The import dispatch counter at detection time.</summary>
public long DispatchIndex { get; }
/// <summary>The first two guest ABI arguments at stall detection.</summary>
public ulong Argument0 { get; }
public ulong Argument1 { get; }
/// <summary>The resolved HLE export, when the NID is registered.</summary>
public string? LibraryName { get; }
public string? FunctionName { get; }
public bool IsResolved => !string.IsNullOrWhiteSpace(FunctionName);
/// <summary>A human-readable one-line summary of the stall.</summary>
public string Detail { get; }
}
@@ -0,0 +1,66 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using SharpEmu.HLE;
namespace SharpEmu.Core.Cpu.Debugging;
/// <summary>
/// A live view of the guest CPU state at a dispatch boundary, handed to an
/// <see cref="ICpuDebugHook"/> so a debugger can read and mutate registers and
/// guest memory without taking a dependency on the concrete
/// <c>CpuContext</c>/<c>CpuDispatcher</c> types.
/// </summary>
/// <remarks>
/// The frame instance is only valid for the duration of the hook call that
/// receives it (between <see cref="ICpuDebugHook.OnFrameEnter"/> and the
/// matching <see cref="ICpuDebugHook.OnFrameExit"/>). Reads and writes are
/// forwarded straight to the underlying guest context, so mutations made from
/// a hook are observed by the CPU backend when it resumes the frame.
/// </remarks>
public interface ICpuDebugFrame
{
/// <summary>The kind of frame being executed.</summary>
CpuDebugFrameKind Kind { get; }
/// <summary>The guest ABI generation this frame targets.</summary>
Generation Generation { get; }
/// <summary>The guest virtual address the frame begins executing at.</summary>
ulong EntryPoint { get; }
/// <summary>
/// A human-readable label for the frame (process image name or module name).
/// </summary>
string Label { get; }
/// <summary>Guest-addressable memory for this frame.</summary>
ICpuMemory Memory { get; }
/// <summary>Reads a general-purpose register.</summary>
ulong GetRegister(CpuRegister register);
/// <summary>Overwrites a general-purpose register.</summary>
void SetRegister(CpuRegister register, ulong value);
/// <summary>The instruction pointer.</summary>
ulong Rip { get; set; }
/// <summary>The flags register.</summary>
ulong Rflags { get; set; }
/// <summary>The FS segment base (guest TLS pointer).</summary>
ulong FsBase { get; }
/// <summary>The GS segment base.</summary>
ulong GsBase { get; }
/// <summary>Reads the 128-bit value of an XMM register.</summary>
void GetXmm(int registerIndex, out ulong low, out ulong high);
/// <summary>
/// The import stubs (guest address to NID) resolved for this frame, so a
/// debugger can annotate calls into HLE exports.
/// </summary>
IReadOnlyDictionary<ulong, string> ImportStubs { get; }
}
@@ -0,0 +1,49 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using SharpEmu.HLE;
namespace SharpEmu.Core.Cpu.Debugging;
/// <summary>
/// The seam the CPU dispatcher uses to notify an attached debugger when guest
/// execution crosses a frame boundary. Implemented outside of Core (for
/// example by <c>SharpEmu.Debugger</c>) and supplied through
/// <see cref="CpuExecutionOptions.DebugHook"/>.
/// </summary>
/// <remarks>
/// This is intentionally coarse-grained: it exposes the entry and exit of each
/// dispatched frame rather than per-instruction stepping. Per-instruction
/// control requires cooperation from the native execution backend and is layered
/// on top of this seam as the backend gains support; keeping the dispatcher-level
/// contract stable lets the debugger infrastructure exist independently of that
/// work. Implementations must be thread-safe: frames may be dispatched from the
/// dedicated emulation thread while a debug server services clients on its own
/// threads.
/// </remarks>
public interface ICpuDebugHook
{
/// <summary>
/// Invoked immediately before the native backend begins executing a frame.
/// The debugger may inspect or mutate <paramref name="frame"/> and may block
/// the calling thread (for example, to honour a pause request) before
/// returning to allow execution to proceed.
/// </summary>
void OnFrameEnter(ICpuDebugFrame frame);
/// <summary>
/// Invoked after a frame completes, whether it returned to the host or
/// terminated with an error. <paramref name="frame"/> reflects the final
/// guest state.
/// </summary>
void OnFrameExit(ICpuDebugFrame frame, OrbisGen2Result result);
/// <summary>
/// Invoked from the emulation thread when the backend detects an execution
/// stall (for example a mutex spin loop) in the running frame, before it
/// forces the guest out of the loop. As with <see cref="OnFrameEnter"/>, the
/// implementation may inspect <paramref name="frame"/> and block to honour a
/// break before returning to let the backend proceed.
/// </summary>
void OnStall(ICpuDebugFrame frame, CpuStallInfo info);
}
@@ -0,0 +1,305 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using System.Numerics;
namespace SharpEmu.Core.Cpu.Emulation;
/// <summary>
/// Pure software implementations of the BMI1, BMI2 and ABM general-purpose-register
/// bit-manipulation instructions.
///
/// The direct-execution backend runs guest PS5 code natively on the host CPU. The PS5's
/// Zen 2 cores implement BMI1/BMI2/ABM, but a host CPU that predates those extensions raises
/// #UD (STATUS_ILLEGAL_INSTRUCTION) when it meets one of these opcodes. This class provides the
/// register-only arithmetic so the exception handler can finish the instruction in software and
/// resume, instead of aborting the title.
///
/// The methods deliberately operate on plain integers rather than on the OS CONTEXT record so the
/// semantics can be unit-tested in isolation; the unsafe register/memory plumbing lives in the
/// backend adapter. Each method returns the result already masked to the operand width and, where
/// the instruction is defined to touch flags, updates <paramref name="eflags"/> in place. Flags
/// documented as "undefined" by the vendor manuals are left untouched so behaviour stays
/// deterministic across hosts.
/// </summary>
public static class BmiInstructionEmulator
{
private const uint FlagCarry = 1u << 0;
private const uint FlagZero = 1u << 6;
private const uint FlagSign = 1u << 7;
private const uint FlagOverflow = 1u << 11;
private static ulong WidthMask(GprOperandSize size) =>
size == GprOperandSize.Bits64 ? ulong.MaxValue : 0xFFFF_FFFFUL;
private static int WidthBits(GprOperandSize size) => (int)size;
private static bool SignSet(ulong value, GprOperandSize size) =>
size == GprOperandSize.Bits64 ? (value >> 63) != 0 : (value & 0x8000_0000UL) != 0;
private static uint WithFlag(uint eflags, uint flag, bool set) =>
set ? eflags | flag : eflags & ~flag;
// Applies the CF/ZF/SF/OF set shared by ANDN/BLS*/BZHI: OF is always cleared, ZF and SF follow
// the result, and the caller supplies CF because each instruction defines it differently.
private static uint ApplyLogicFlags(uint eflags, ulong result, GprOperandSize size, bool carry)
{
eflags = WithFlag(eflags, FlagCarry, carry);
eflags = WithFlag(eflags, FlagZero, result == 0);
eflags = WithFlag(eflags, FlagSign, SignSet(result, size));
eflags = WithFlag(eflags, FlagOverflow, false);
return eflags;
}
/// <summary>ANDN: <c>dest = (~src1) &amp; src2</c>. CF and OF are cleared.</summary>
public static ulong Andn(ulong src1, ulong src2, GprOperandSize size, ref uint eflags)
{
var result = (~src1 & src2) & WidthMask(size);
eflags = ApplyLogicFlags(eflags, result, size, carry: false);
return result;
}
/// <summary>BLSI: isolate the lowest set bit, <c>dest = (-src) &amp; src</c>. CF = (src != 0).</summary>
public static ulong Blsi(ulong src, GprOperandSize size, ref uint eflags)
{
var mask = WidthMask(size);
var s = src & mask;
var result = ((0UL - s) & s) & mask;
eflags = ApplyLogicFlags(eflags, result, size, carry: s != 0);
return result;
}
/// <summary>BLSMSK: mask up to and including the lowest set bit, <c>dest = (src - 1) ^ src</c>. CF = (src == 0).</summary>
public static ulong Blsmsk(ulong src, GprOperandSize size, ref uint eflags)
{
var mask = WidthMask(size);
var s = src & mask;
var result = ((s - 1) ^ s) & mask;
eflags = ApplyLogicFlags(eflags, result, size, carry: s == 0);
return result;
}
/// <summary>BLSR: reset the lowest set bit, <c>dest = (src - 1) &amp; src</c>. CF = (src == 0).</summary>
public static ulong Blsr(ulong src, GprOperandSize size, ref uint eflags)
{
var mask = WidthMask(size);
var s = src & mask;
var result = ((s - 1) & s) & mask;
eflags = ApplyLogicFlags(eflags, result, size, carry: s == 0);
return result;
}
/// <summary>
/// BEXTR: extract <c>len</c> bits of <paramref name="src"/> starting at bit <c>start</c>, where
/// start = control[7:0] and len = control[15:8]. Only ZF (per result) and cleared CF/OF are defined.
/// </summary>
public static ulong Bextr(ulong src, ulong control, GprOperandSize size, ref uint eflags)
{
var bits = WidthBits(size);
var start = (int)(control & 0xFF);
var length = (int)((control >> 8) & 0xFF);
ulong result;
if (start >= bits)
{
result = 0;
}
else
{
var shifted = (src & WidthMask(size)) >> start;
if (length == 0)
{
result = 0;
}
else if (length >= bits)
{
result = shifted;
}
else
{
result = shifted & ((1UL << length) - 1);
}
}
result &= WidthMask(size);
eflags = WithFlag(eflags, FlagZero, result == 0);
eflags = WithFlag(eflags, FlagCarry, false);
eflags = WithFlag(eflags, FlagOverflow, false);
return result;
}
/// <summary>
/// BZHI: zero the bits of <paramref name="src"/> from bit position <c>index[7:0]</c> upward.
/// CF is set when the requested position is at or beyond the operand width.
/// </summary>
public static ulong Bzhi(ulong src, ulong index, GprOperandSize size, ref uint eflags)
{
var bits = WidthBits(size);
var mask = WidthMask(size);
var s = src & mask;
var n = (int)(index & 0xFF);
ulong result;
bool carry;
if (n >= bits)
{
result = s;
carry = true;
}
else
{
result = s & ((1UL << n) - 1);
carry = false;
}
eflags = ApplyLogicFlags(eflags, result, size, carry);
return result;
}
/// <summary>TZCNT: count trailing zero bits. If src == 0 the result is the operand width and CF is set.</summary>
public static ulong Tzcnt(ulong src, GprOperandSize size, ref uint eflags)
{
var bits = WidthBits(size);
var s = src & WidthMask(size);
ulong result;
bool carry;
if (s == 0)
{
result = (ulong)bits;
carry = true;
}
else
{
result = (ulong)(size == GprOperandSize.Bits64
? BitOperations.TrailingZeroCount(s)
: BitOperations.TrailingZeroCount((uint)s));
carry = false;
}
eflags = WithFlag(eflags, FlagCarry, carry);
eflags = WithFlag(eflags, FlagZero, result == 0);
return result;
}
/// <summary>LZCNT: count leading zero bits. If src == 0 the result is the operand width and CF is set.</summary>
public static ulong Lzcnt(ulong src, GprOperandSize size, ref uint eflags)
{
var bits = WidthBits(size);
var s = src & WidthMask(size);
ulong result;
bool carry;
if (s == 0)
{
result = (ulong)bits;
carry = true;
}
else
{
result = (ulong)(size == GprOperandSize.Bits64
? BitOperations.LeadingZeroCount(s)
: BitOperations.LeadingZeroCount((uint)s));
carry = false;
}
eflags = WithFlag(eflags, FlagCarry, carry);
eflags = WithFlag(eflags, FlagZero, result == 0);
return result;
}
/// <summary>RORX: rotate <paramref name="src"/> right by <paramref name="count"/> (masked to the operand width). No flags.</summary>
public static ulong Rorx(ulong src, int count, GprOperandSize size)
{
var bits = WidthBits(size);
var mask = WidthMask(size);
var s = src & mask;
var rotate = count & (bits - 1);
if (rotate == 0)
{
return s;
}
return ((s >> rotate) | (s << (bits - rotate))) & mask;
}
/// <summary>SARX: arithmetic shift right by <paramref name="count"/> (masked to the operand width). No flags.</summary>
public static ulong Sarx(ulong src, int count, GprOperandSize size)
{
var bits = WidthBits(size);
var mask = WidthMask(size);
var shift = count & (bits - 1);
if (size == GprOperandSize.Bits64)
{
return (ulong)((long)src >> shift);
}
return (ulong)(uint)((int)(uint)src >> shift) & mask;
}
/// <summary>SHLX: logical shift left by <paramref name="count"/> (masked to the operand width). No flags.</summary>
public static ulong Shlx(ulong src, int count, GprOperandSize size)
{
var bits = WidthBits(size);
var mask = WidthMask(size);
var shift = count & (bits - 1);
return (src << shift) & mask;
}
/// <summary>SHRX: logical shift right by <paramref name="count"/> (masked to the operand width). No flags.</summary>
public static ulong Shrx(ulong src, int count, GprOperandSize size)
{
var bits = WidthBits(size);
var mask = WidthMask(size);
var s = src & mask;
var shift = count & (bits - 1);
return s >> shift;
}
/// <summary>PDEP: deposit contiguous low bits of <paramref name="src"/> into the positions selected by <paramref name="mask"/>. No flags.</summary>
public static ulong Pdep(ulong src, ulong mask, GprOperandSize size)
{
var bits = WidthBits(size);
var selector = mask & WidthMask(size);
ulong result = 0;
var bit = 0;
for (var i = 0; i < bits; i++)
{
var position = 1UL << i;
if ((selector & position) != 0)
{
if (((src >> bit) & 1UL) != 0)
{
result |= position;
}
bit++;
}
}
return result & WidthMask(size);
}
/// <summary>PEXT: gather the bits of <paramref name="src"/> selected by <paramref name="mask"/> into contiguous low bits. No flags.</summary>
public static ulong Pext(ulong src, ulong mask, GprOperandSize size)
{
var bits = WidthBits(size);
var selector = mask & WidthMask(size);
ulong result = 0;
var bit = 0;
for (var i = 0; i < bits; i++)
{
if ((selector & (1UL << i)) != 0)
{
if (((src >> i) & 1UL) != 0)
{
result |= 1UL << bit;
}
bit++;
}
}
return result & WidthMask(size);
}
}
@@ -0,0 +1,14 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
namespace SharpEmu.Core.Cpu.Emulation;
/// <summary>
/// Operand width for the emulated general-purpose-register instructions. The numeric value is the
/// bit width, so it can double as the "count leading/trailing zeros of an all-zero source" result.
/// </summary>
public enum GprOperandSize
{
Bits32 = 32,
Bits64 = 64,
}
@@ -0,0 +1,71 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
namespace SharpEmu.Core.Cpu.Emulation;
/// <summary>
/// Pure software implementation of the bit-field math behind AMD's SSE4a EXTRQ/INSERTQ
/// (immediate-form) instructions.
///
/// The direct-execution backend runs guest PS5 code natively on the host CPU. The PS5's Zen 2
/// cores implement AMD-only SSE4a (EXTRQ/INSERTQ), but Intel hosts - and Rosetta 2 on Apple
/// Silicon - do not, so they raise #UD (STATUS_ILLEGAL_INSTRUCTION) instead of executing the
/// opcode. SharpEmu already rewrites one specific compiled EXTRQ+VPBLENDD idiom at load time
/// (see <see cref="Native.Sse4aExtrqBlendPatch"/>), but any other occurrence of EXTRQ/INSERTQ -
/// a different register allocation, a title built with a different compiler version, and so on
/// - still aborts the title. This class ported from Kyty's
/// <c>Loader::X64InstructionEmulator::TryEmulateSse4a</c> provides the general bit-field
/// extract/insert so the illegal-instruction handler can finish *any* immediate-form
/// EXTRQ/INSERTQ in software and resume, instead of relying on a single hard-coded byte pattern.
///
/// The methods operate on plain 64-bit integers rather than the OS CONTEXT record so the bit
/// math can be unit-tested in isolation; the unsafe CONTEXT/XMM plumbing lives in the backend
/// adapter (<see cref="Native.DirectExecutionBackend"/>).
/// </summary>
public static class Sse4aBitFieldEmulator
{
public static bool IsValidBitField(int length, int index)
{
var len = length & 0x3F;
var idx = index & 0x3F;
return (len != 0 || idx == 0) && (len == 0 ? idx == 0 : idx + len <= 64);
}
public static ulong ExtractBitField(ulong value, int length, int index)
{
var len = length & 0x3F;
var idx = index & 0x3F;
if (!IsValidBitField(length, index))
{
return 0;
}
if (len == 0)
{
return value;
}
var mask = len == 64 ? ulong.MaxValue : (1UL << len) - 1;
return (value >> idx) & mask;
}
public static ulong InsertBitField(ulong destination, ulong source, int length, int index)
{
var len = length & 0x3F;
var idx = index & 0x3F;
if (!IsValidBitField(length, index))
{
return destination;
}
if (len == 0)
{
return source;
}
var fieldMask = len == 64 ? ulong.MaxValue : (1UL << len) - 1;
var destinationClearMask = fieldMask << idx;
var sourceField = (source & fieldMask) << idx;
return (destination & ~destinationClearMask) | sourceField;
}
}
@@ -0,0 +1,199 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using System.Threading;
using Iced.Intel;
using SharpEmu.Core.Cpu.Emulation;
namespace SharpEmu.Core.Cpu.Native;
// General software fallback for the AMD-only instructions PS5 titles occasionally emit that a
// Zen 2-only host implements but Intel hosts (and Rosetta 2 on Apple Silicon) do not:
// - SSE4a EXTRQ/INSERTQ, immediate form
// - MONITORX/MWAITX
//
// This is a direct port of Kyty's Loader::X64InstructionEmulator (TryEmulateSse4a /
// TryEmulateMonitorxMwaitx). SharpEmu already special-cases exactly one compiled EXTRQ+VPBLENDD
// byte sequence at load time (Sse4aExtrqBlendPatch), which only helps the one idiom it was
// reverse-engineered from. This file is a general, fault-time fallback that engages for any
// immediate-form EXTRQ/INSERTQ or MONITORX/MWAITX the narrower patch (or a title using a
// different compiler/register allocation) does not cover, complementing rather than replacing
// it: the load-time patch still avoids paying the fault-and-recover cost on the hot path it was
// built for, while this method is the safety net for everything else.
//
// This is deliberately additive: DirectExecutionBackend.IllegalInstruction.cs (the BMI1/BMI2/ABM
// fallback) is untouched, and this method is only reached from VectoredHandler after that one
// has already declined to handle the fault.
public sealed partial class DirectExecutionBackend
{
// Byte offset of Xmm0 within the Win64 CONTEXT record: FltSave (the XMM_SAVE_AREA32/FXSAVE
// image) starts right after Rip at offset 256, and XmmRegisters[0] sits 160 bytes into that
// area (32-byte header + 8 legacy x87/MMX slots x 16 bytes). 256 + 160 = 416 (0x1A0). Cross-
// checked against this file's own Win64ContextSize (0x4D0): rebuilding the whole CONTEXT
// layout field-by-field from offset 0 lands on the same 0x4D0 total, which would not happen
// if this offset (or anything before it) were wrong.
private const int Win64ContextXmm0Offset = 0x1A0;
private static int _sse4aSoftwareFallbackAnnounced;
private static long _sse4aInstructionsEmulated;
private static int _monitorxSoftwareFallbackAnnounced;
private static long _monitorxInstructionsEmulated;
private unsafe bool TryRecoverAmdCompatInstruction(void* contextRecord, ulong rip)
{
if (TryRecoverMonitorxMwaitx(contextRecord, rip))
{
return true;
}
// MONITORX/MWAITX above only ever reads guest code memory and rewrites RIP, both of
// which the POSIX signal bridge (DirectExecutionBackend.PosixSignals.cs) faithfully
// round-trips through the real ucontext, so it works on every supported OS. EXTRQ/
// INSERTQ additionally read and write an XMM register: on Windows contextRecord is the
// live CONTEXT the OS resumes the thread from, so touching the Xmm0.. slots is visible
// to the guest, and on Linux the bridge copies the mcontext's FXSAVE image into the
// Xmm0.. slots and writes them back through sigreturn (_posixXmmContextBridged). On
// Darwin the XMM area is still a zeroed scratch buffer - running this there would
// silently compute a result from stale bytes and then discard whatever it "wrote", so
// the recovery declines until that bridge exists.
return (OperatingSystem.IsWindows() || _posixXmmContextBridged) &&
TryRecoverSse4aExtractInsert(contextRecord, rip);
}
private unsafe bool TryRecoverMonitorxMwaitx(void* contextRecord, ulong rip)
{
// MONITORX (0F 01 FA) and MWAITX (0F 01 FB) are fixed 3-byte encodings with no
// ModRM/SIB/displacement/immediate, so a raw byte compare is sufficient and unambiguous.
var opcode = new byte[3];
if (!TryReadHostBytes(rip, opcode) ||
opcode[0] != 0x0F || opcode[1] != 0x01 || (opcode[2] != 0xFA && opcode[2] != 0xFB))
{
return false;
}
// PS5 titles use this pair in idle/wait loops: MONITORX arms a monitor on a cache line
// and MWAITX blocks until that line is written (or a timeout elapses). Hosts without
// the extension raise #UD on either one. We do not model the monitor itself, only its
// observable effect on guest forward progress: MONITORX becomes a no-op (arming a
// watch we never honour has no side effect of its own) and MWAITX becomes a plain
// thread yield, i.e. treat the awaited condition as already satisfied so the guest
// loop keeps making progress instead of executing an illegal opcode forever.
if (opcode[2] == 0xFB)
{
Thread.Yield();
}
WriteCtxU64(contextRecord, CTX_RIP, rip + 3);
Interlocked.Increment(ref _monitorxInstructionsEmulated);
if (Interlocked.Exchange(ref _monitorxSoftwareFallbackAnnounced, 1) == 0)
{
Console.Error.WriteLine(
"[LOADER][INFO] Host lacks AMD MONITORX/MWAITX used by the guest; " +
"emulating those instructions in software.");
}
return true;
}
private unsafe bool TryRecoverSse4aExtractInsert(void* contextRecord, ulong rip)
{
if (!OperatingSystem.IsWindows() && !_posixXmmContextBridged ||
!TryReadFaultingInstruction(rip, out var instruction))
{
return false;
}
var isExtrq = instruction.Mnemonic == Mnemonic.Extrq;
var isInsertq = instruction.Mnemonic == Mnemonic.Insertq;
if (!isExtrq && !isInsertq)
{
return false;
}
if (isExtrq && instruction.OpCount != 3 || isInsertq && instruction.OpCount != 4)
{
return false;
}
if (instruction.GetOpKind(0) != OpKind.Register ||
!TryGetXmmOffset(instruction.GetOpRegister(0), out var destOffset))
{
return false;
}
var destLow = ReadCtxU64(contextRecord, destOffset);
if (isExtrq)
{
var length = (int)instruction.GetImmediate(1);
var index = (int)instruction.GetImmediate(2);
if (!Sse4aBitFieldEmulator.IsValidBitField(length, index))
{
return false;
}
WriteCtxU64(contextRecord, destOffset, Sse4aBitFieldEmulator.ExtractBitField(destLow, length, index));
WriteCtxU64(contextRecord, destOffset + 8, 0);
}
else
{
if (instruction.GetOpKind(1) != OpKind.Register ||
!TryGetXmmOffset(instruction.GetOpRegister(1), out var srcOffset))
{
return false;
}
var length = (int)instruction.GetImmediate(2);
var index = (int)instruction.GetImmediate(3);
if (!Sse4aBitFieldEmulator.IsValidBitField(length, index))
{
return false;
}
WriteCtxU64(contextRecord, destOffset, Sse4aBitFieldEmulator.InsertBitField(
destLow, ReadCtxU64(contextRecord, srcOffset), length, index));
WriteCtxU64(contextRecord, destOffset + 8, 0);
}
WriteCtxU64(contextRecord, CTX_RIP, rip + (ulong)instruction.Length);
Interlocked.Increment(ref _sse4aInstructionsEmulated);
if (Interlocked.Exchange(ref _sse4aSoftwareFallbackAnnounced, 1) == 0)
{
Console.Error.WriteLine(
"[LOADER][INFO] Host lacks SSE4a EXTRQ/INSERTQ used by the guest; " +
"emulating those instructions in software.");
}
return true;
}
// Maps an Iced XMM register to its byte offset in the Win64 CONTEXT record. Written as an
// explicit switch (rather than arithmetic on the Register enum) to match the style already
// used by TryGetGprSlot/TryGetGpr64Offset in DirectExecutionBackend.IllegalInstruction.cs.
private static bool TryGetXmmOffset(Register register, out int offset)
{
switch (register)
{
case Register.XMM0: offset = Win64ContextXmm0Offset + 16 * 0; return true;
case Register.XMM1: offset = Win64ContextXmm0Offset + 16 * 1; return true;
case Register.XMM2: offset = Win64ContextXmm0Offset + 16 * 2; return true;
case Register.XMM3: offset = Win64ContextXmm0Offset + 16 * 3; return true;
case Register.XMM4: offset = Win64ContextXmm0Offset + 16 * 4; return true;
case Register.XMM5: offset = Win64ContextXmm0Offset + 16 * 5; return true;
case Register.XMM6: offset = Win64ContextXmm0Offset + 16 * 6; return true;
case Register.XMM7: offset = Win64ContextXmm0Offset + 16 * 7; return true;
case Register.XMM8: offset = Win64ContextXmm0Offset + 16 * 8; return true;
case Register.XMM9: offset = Win64ContextXmm0Offset + 16 * 9; return true;
case Register.XMM10: offset = Win64ContextXmm0Offset + 16 * 10; return true;
case Register.XMM11: offset = Win64ContextXmm0Offset + 16 * 11; return true;
case Register.XMM12: offset = Win64ContextXmm0Offset + 16 * 12; return true;
case Register.XMM13: offset = Win64ContextXmm0Offset + 16 * 13; return true;
case Register.XMM14: offset = Win64ContextXmm0Offset + 16 * 14; return true;
case Register.XMM15: offset = Win64ContextXmm0Offset + 16 * 15; return true;
default:
offset = 0;
return false;
}
}
}
@@ -7,6 +7,7 @@ using System.Collections.Generic;
using System.Linq;
using System.Runtime.InteropServices;
using System.Threading;
using SharpEmu.Core.Cpu.Disasm;
using SharpEmu.HLE;
using SharpEmu.Logging;
@@ -16,6 +17,55 @@ public sealed partial class DirectExecutionBackend
{
private static readonly ConcurrentDictionary<ulong, byte> _knownExecutablePages = new();
private static readonly bool _perfHleHistogram =
string.Equals(System.Environment.GetEnvironmentVariable("SHARPEMU_PERF_HLE"), "1", System.StringComparison.Ordinal);
private static readonly System.Collections.Concurrent.ConcurrentDictionary<string, long> _perfHleCounts = new();
private static long _perfHleTotal;
private static long _perfHleDispatchTicks;
private static void RecordPerfHleDispatchTime(long ticks)
{
var total = System.Threading.Interlocked.Add(ref _perfHleDispatchTicks, ticks);
var calls = System.Threading.Interlocked.Read(ref _perfHleTotal);
if (calls > 0 && calls % 500000 == 0)
{
var avgUs = (double)total / System.Diagnostics.Stopwatch.Frequency * 1_000_000.0 / calls;
System.Console.Error.WriteLine($"[PERF][HLE] managed_dispatch_avg={avgUs:F3}us total_managed_s={(double)total / System.Diagnostics.Stopwatch.Frequency:F2}");
}
}
private static readonly bool _perfHleNoDict =
string.Equals(System.Environment.GetEnvironmentVariable("SHARPEMU_PERF_HLE_NODICT"), "1", System.StringComparison.Ordinal);
private static void RecordPerfHleCall(string name)
{
var total = System.Threading.Interlocked.Increment(ref _perfHleTotal);
if (!_perfHleNoDict)
{
_perfHleCounts.AddOrUpdate(name, 1, static (_, v) => v + 1);
}
if (total % 500000 == 0 && !_perfHleNoDict)
{
// Snapshot via foreach (a safe moving enumerator) before sorting.
// LINQ over a ConcurrentDictionary uses ICollection.CopyTo, which
// throws ArgumentException if another thread adds a key between the
// Count read and the copy — that exception was being swallowed into
// a CPU_TRAP return and crashing the guest.
var snapshot = new System.Collections.Generic.List<System.Collections.Generic.KeyValuePair<string, long>>(_perfHleCounts.Count + 16);
foreach (var kvp in _perfHleCounts)
{
snapshot.Add(kvp);
}
var top = snapshot
.OrderByDescending(kvp => kvp.Value)
.Take(20)
.Select(kvp => $"{kvp.Key}={kvp.Value}");
System.Console.Error.WriteLine($"[PERF][HLE] total={total} top: {string.Join(", ", top)}");
}
}
private void RecordRecentImportTrace(
long dispatchIndex,
string nid,
@@ -24,101 +74,40 @@ public sealed partial class DirectExecutionBackend
ulong arg1,
ulong arg2)
{
_recentImportTrace[_recentImportTraceWriteIndex] = new RecentImportTraceEntry(
var trace = _recentImportTrace;
trace[_recentImportTraceWriteIndex] = new RecentImportTraceEntry(
dispatchIndex,
nid,
returnRip,
arg0,
arg1,
arg2);
_recentImportTraceWriteIndex = (_recentImportTraceWriteIndex + 1) % _recentImportTrace.Length;
if (_recentImportTraceCount < _recentImportTrace.Length)
arg2,
GuestThreadExecution.CurrentGuestThreadHandle,
Environment.CurrentManagedThreadId);
_recentImportTraceWriteIndex = (_recentImportTraceWriteIndex + 1) % trace.Length;
if (_recentImportTraceCount < trace.Length)
{
_recentImportTraceCount++;
}
}
private void RecordDeferredBootstrapTrace(
long dispatchIndex,
ulong op,
ulong symbolPointer,
ulong outputPointer,
ulong returnRip)
{
lock (_deferredBootstrapTraceGate)
{
_deferredBootstrapTrace[_deferredBootstrapTraceWriteIndex] = new DeferredBootstrapTraceEntry(
dispatchIndex,
op,
symbolPointer,
outputPointer,
returnRip);
_deferredBootstrapTraceWriteIndex =
(_deferredBootstrapTraceWriteIndex + 1) % _deferredBootstrapTrace.Length;
if (_deferredBootstrapTraceCount < _deferredBootstrapTrace.Length)
{
_deferredBootstrapTraceCount++;
}
}
}
private void DrainDeferredBootstrapTraces()
{
if (!_logBootstrap)
{
return;
}
DeferredBootstrapTraceEntry[] pending;
lock (_deferredBootstrapTraceGate)
{
if (_deferredBootstrapTraceCount == 0)
{
return;
}
pending = new DeferredBootstrapTraceEntry[_deferredBootstrapTraceCount];
var readIndex = (_deferredBootstrapTraceWriteIndex - _deferredBootstrapTraceCount +
_deferredBootstrapTrace.Length) % _deferredBootstrapTrace.Length;
for (var i = 0; i < _deferredBootstrapTraceCount; i++)
{
pending[i] = _deferredBootstrapTrace[(readIndex + i) % _deferredBootstrapTrace.Length];
}
_deferredBootstrapTraceCount = 0;
}
foreach (var entry in pending)
{
var symbolText = "<unreadable>";
if (TryReadAsciiZ(entry.SymbolPointer, 256, out var sym))
{
symbolText = sym;
}
Console.Error.WriteLine(
$"[LOADER][TRACE] bootstrap_call#{entry.DispatchIndex}: op=0x{entry.Op:X16} " +
$"sym_ptr=0x{entry.SymbolPointer:X16} sym='{symbolText}' " +
$"out_ptr=0x{entry.OutputPointer:X16} ret=0x{entry.ReturnRip:X16}");
}
}
private void DumpRecentImportTrace()
{
if (_recentImportTraceCount == 0)
var trace = _recentImportTrace;
if (trace is null || _recentImportTraceCount == 0)
{
return;
}
Log.Info($" Recent import calls ({_recentImportTraceCount}):");
int num = (_recentImportTraceWriteIndex - _recentImportTraceCount + _recentImportTrace.Length) % _recentImportTrace.Length;
Log.Info($" Recent import calls for managed={Environment.CurrentManagedThreadId} guest=0x{GuestThreadExecution.CurrentGuestThreadHandle:X16} ({_recentImportTraceCount}):");
int num = (_recentImportTraceWriteIndex - _recentImportTraceCount + trace.Length) % trace.Length;
for (int i = 0; i < _recentImportTraceCount; i++)
{
int num2 = (num + i) % _recentImportTrace.Length;
var entry = _recentImportTrace[num2];
int num2 = (num + i) % trace.Length;
var entry = trace[num2];
if (!string.IsNullOrEmpty(entry.Nid))
{
Log.Info(
$" #{entry.DispatchIndex} nid={entry.Nid} ret=0x{entry.ReturnRip:X16} " +
$" #{entry.DispatchIndex} managed={entry.ManagedThreadId} guest=0x{entry.GuestThreadHandle:X16} nid={entry.Nid} ret=0x{entry.ReturnRip:X16} " +
$"rdi=0x{entry.Arg0:X16} rsi=0x{entry.Arg1:X16} rdx=0x{entry.Arg2:X16}");
}
}
@@ -184,6 +173,50 @@ public sealed partial class DirectExecutionBackend
{
return;
}
const int preludeSize = 192;
Span<byte> prelude = stackalloc byte[preludeSize];
if (returnRip >= preludeSize && cpuContext.Memory.TryRead(returnRip - preludeSize, prelude))
{
Console.Error.WriteLine(
$"[LOADER][TRACE] Import#{dispatchIndex} pre-return bytes @0x{returnRip - preludeSize:X16}: " +
BitConverter.ToString(prelude.ToArray()).Replace("-", " "));
List<DecodedInst>? bestCallChain = null;
var preludeAddress = returnRip - preludeSize;
for (var startOffset = 0; startOffset < preludeSize; startOffset++)
{
var cursor = preludeAddress + (ulong)startOffset;
var candidate = new List<DecodedInst>();
while (cursor < returnRip && candidate.Count < 96 &&
IcedDecoder.TryReadGuestBytes(cpuContext.Memory, cursor, 15, out var instructionBytes) &&
IcedDecoder.TryDecode(cursor, instructionBytes, out var instruction) &&
instruction.Length > 0 &&
cursor + (ulong)instruction.Length <= returnRip)
{
candidate.Add(instruction);
cursor += (ulong)instruction.Length;
}
if (cursor == returnRip &&
candidate.Count > 0 &&
string.Equals(candidate[^1].Mnemonic, "Call", StringComparison.OrdinalIgnoreCase) &&
(bestCallChain is null || candidate.Count > bestCallChain.Count))
{
bestCallChain = candidate;
}
}
if (bestCallChain is not null)
{
Console.Error.WriteLine($"[LOADER][TRACE] Import#{dispatchIndex} pre-return disassembly:");
foreach (var instruction in bestCallChain.TakeLast(32))
{
Console.Error.WriteLine(
$"[LOADER][TRACE] 0x{instruction.Rip:X16}: {instruction.Text} " +
$"bytes={IcedDecoder.FormatBytes(instruction.Bytes)}");
}
}
}
Span<byte> destination = stackalloc byte[128];
if (!cpuContext.Memory.TryRead(returnRip, destination))
{
@@ -235,15 +268,14 @@ public sealed partial class DirectExecutionBackend
ulong callRip = returnRip + (ulong)i;
ulong target = unchecked((ulong)((long)(callRip + 5) + rel32));
Log.Debug($"Import#{dispatchIndex} near-call @{callRip:X16}: target=0x{target:X16}");
var importEntries = _importEntries;
for (int importIndex = 0; importIndex < importEntries.Length; importIndex++)
for (int importIndex = 0; importIndex < _importEntries.Length; importIndex++)
{
if (importEntries[importIndex].Address != target)
if (_importEntries[importIndex].Address != target)
{
continue;
}
string nid = importEntries[importIndex].Nid;
string nid = _importEntries[importIndex].Nid;
if (_moduleManager.TryGetExport(nid, out var export))
{
Log.Debug(
@@ -270,14 +302,14 @@ public sealed partial class DirectExecutionBackend
{
Log.Debug(
$"Import#{dispatchIndex} near-call PLT slot: [0x{slot:X16}] = 0x{slotTarget:X16}");
for (int importIndex = 0; importIndex < importEntries.Length; importIndex++)
for (int importIndex = 0; importIndex < _importEntries.Length; importIndex++)
{
if (importEntries[importIndex].Address != slotTarget)
if (_importEntries[importIndex].Address != slotTarget)
{
continue;
}
string nid = importEntries[importIndex].Nid;
string nid = _importEntries[importIndex].Nid;
if (_moduleManager.TryGetExport(nid, out var export))
{
Log.Debug(
@@ -301,6 +333,28 @@ public sealed partial class DirectExecutionBackend
return value == 65534 || value == 4294967294u || value == 18446744073709551614uL;
}
private static ulong ParseOptionalHexAddress(string? value)
{
if (string.IsNullOrWhiteSpace(value))
{
return 0;
}
var text = value.Trim();
if (text.StartsWith("0x", StringComparison.OrdinalIgnoreCase))
{
text = text[2..];
}
return ulong.TryParse(
text,
System.Globalization.NumberStyles.HexNumber,
System.Globalization.CultureInfo.InvariantCulture,
out var address)
? address
: 0;
}
private static bool IsPlausibleReturnAddress(ulong address)
{
return address >= 12884901888L && address < 17592186044416L && !IsUnresolvedSentinel(address);
@@ -17,9 +17,17 @@ public sealed partial class DirectExecutionBackend
{
private const ulong LazyCommitWindowBytes = 0x0200_0000UL;
private static int _lazyCommitTraceCount;
private static int _guestAllocatorHoleRecoveries;
private static int _auxiliaryThreadExecuteFaultRecoveries;
private unsafe void SetupExceptionHandler()
{
if (!OperatingSystem.IsWindows())
{
SetupPosixExceptionHandler();
return;
}
if (!string.Equals(Environment.GetEnvironmentVariable("SHARPEMU_DISABLE_RAW_HANDLER"), "1", StringComparison.Ordinal))
{
_rawExceptionHandlerStub = CreateExceptionHandlerTrampoline(RawVectoredHandlerPtrManaged);
@@ -102,20 +110,34 @@ public sealed partial class DirectExecutionBackend
ulong rip = ReadCtxU64(contextRecord, 248);
ulong rsp = ReadCtxU64(contextRecord, 152);
// Thread-mode probe: a hardware exception raised while this thread is inside
// the managed import gateway means the VEH->managed reentry happened from
// cooperative GC mode — a ReversePInvokeBadTransition candidate.
if (LogThreadMode && _threadModeGatewayDepth > 0)
if (TryRecoverGuestInt41(exceptionCode, contextRecord, rip))
{
TraceThreadMode(
$"veh_in_gateway code=0x{exceptionCode:X8} rip=0x{rip:X16} gateway_depth={_threadModeGatewayDepth}");
return -1;
}
if (TryRecoverAuxiliaryThreadExecuteFault(exceptionRecord, contextRecord, rip))
{
return -1;
}
if (exceptionCode == 3221225477u && TryHandleLazyCommittedPage(exceptionRecord, rip, rsp))
{
return -1;
}
if (exceptionCode == 3221225477u &&
TryRecoverGuestAllocatorHole(exceptionRecord, contextRecord, rip))
{
return -1;
}
if (exceptionCode == StatusIllegalInstruction &&
TryRecoverIllegalInstruction(contextRecord, rip))
{
return -1;
}
if (exceptionCode == StatusIllegalInstruction &&
TryRecoverAmdCompatInstruction(contextRecord, rip))
{
return -1;
}
if (IsBenignHostDebugException(exceptionCode))
{
return -1;
@@ -161,6 +183,35 @@ public sealed partial class DirectExecutionBackend
Console.Error.WriteLine($"[LOADER][INFO] Code: 0x{exceptionCode:X8}");
Console.Error.WriteLine($"[LOADER][INFO] Exception Address: 0x{exceptionAddress:X16}");
Console.Error.WriteLine($"[LOADER][INFO] RIP: 0x{rip:X16}");
Console.Error.WriteLine(
$"[LOADER][INFO] Host thread: managed={Environment.CurrentManagedThreadId} " +
$"name='{Thread.CurrentThread.Name ?? "<unnamed>"}'");
if (_activeGuestThreadState is { } activeGuestThread)
{
Console.Error.WriteLine(
$"[LOADER][INFO] Guest thread: handle=0x{activeGuestThread.ThreadHandle:X16} " +
$"name='{activeGuestThread.Name}' state={activeGuestThread.State} " +
$"last_import={activeGuestThread.LastImportNid ?? "<none>"} " +
$"last_ret=0x{activeGuestThread.LastReturnRip:X16}");
Console.Error.WriteLine(
$"[LOADER][INFO] Last import registers: " +
$"rax=0x{Volatile.Read(ref activeGuestThread.LastImportRax):X16} " +
$"result_valid={Volatile.Read(ref activeGuestThread.LastImportResultValid) != 0} " +
$"rdi=0x{activeGuestThread.LastImportRdi:X16} " +
$"rsi=0x{activeGuestThread.LastImportRsi:X16} " +
$"rdx=0x{activeGuestThread.LastImportRdx:X16} " +
$"rcx=0x{activeGuestThread.LastImportRcx:X16} " +
$"r8=0x{activeGuestThread.LastImportR8:X16} " +
$"r9=0x{activeGuestThread.LastImportR9:X16}");
Console.Error.WriteLine(
$"[LOADER][INFO] Last import stack args: " +
$"0=0x{activeGuestThread.LastImportStack0:X16} " +
$"1=0x{activeGuestThread.LastImportStack1:X16} " +
$"2=0x{activeGuestThread.LastImportStack2:X16} " +
$"3=0x{activeGuestThread.LastImportStack3:X16} " +
$"4=0x{activeGuestThread.LastImportStack4:X16} " +
$"5=0x{activeGuestThread.LastImportStack5:X16}");
}
if (TryFormatNearestRuntimeSymbol(rip, out string symbol))
{
Console.Error.WriteLine("[LOADER][INFO] RIP symbol: " + symbol);
@@ -205,20 +256,50 @@ public sealed partial class DirectExecutionBackend
}
try
Console.Error.WriteLine("[LOADER][INFO] Stack qwords (RSP..):");
for (int i = 0; i < 16; i++)
{
Console.Error.WriteLine("[LOADER][INFO] Stack qwords (RSP..):");
for (int i = 0; i < 16; i++)
ulong stackAddr = rsp + (ulong)(i * 8);
if (!TryReadHostQword(stackAddr, out ulong value))
{
ulong stackAddr = rsp + (ulong)(i * 8);
ulong value = (ulong)Marshal.ReadInt64((nint)stackAddr);
Console.Error.WriteLine($"[LOADER][INFO] [rsp+0x{i * 8:X2}] @0x{stackAddr:X16} = 0x{value:X16}");
Console.Error.WriteLine("[LOADER][WARNING] Could not read stack qwords.");
break;
}
Console.Error.WriteLine($"[LOADER][INFO] [rsp+0x{i * 8:X2}] @0x{stackAddr:X16} = 0x{value:X16}");
}
if (string.Equals(
Environment.GetEnvironmentVariable("SHARPEMU_DUMP_FAULT_STACK_WINDOW"),
"1",
StringComparison.Ordinal))
{
Console.Error.WriteLine("[LOADER][INFO] Full fault stack window (RSP-0x300..RSP+0x100):");
var windowStart = rsp >= 0x300 ? rsp - 0x300 : 0;
for (var stackAddr = windowStart; stackAddr < rsp + 0x100; stackAddr += 8)
{
if (!TryReadHostQword(stackAddr, out var value))
{
continue;
}
var relative = unchecked((long)(stackAddr - rsp));
var relativeText = relative >= 0
? $"+0x{relative:X}"
: $"-0x{-relative:X}";
var symbolText = TryFormatNearestRuntimeSymbol(value, out var stackSymbol)
? $" [{stackSymbol}]"
: string.Empty;
Console.Error.WriteLine(
$"[LOADER][INFO] [rsp{relativeText}] " +
$"@0x{stackAddr:X16} = 0x{value:X16}{symbolText}");
}
}
catch
{
Console.Error.WriteLine("[LOADER][WARNING] Could not read stack qwords.");
}
DumpPointerWindow("fault-register-rbx", rbx, 0x60);
DumpPointerWindow("fault-register-rsi", rsi, 0x60);
DumpPointerWindow("fault-register-rdi", rdi, 0x60);
DumpPointerWindow("fault-register-r13", r13, 0x60);
DumpPointerWindow("fault-register-r14", r14, 0x60);
try
{
@@ -226,12 +307,15 @@ public sealed partial class DirectExecutionBackend
ulong frame = rbp;
for (int i = 0; i < 12; i++)
{
if (frame < 140733193388032L || frame > 140737488355327L)
if (frame < 0x10000)
{
break;
}
ulong next = (ulong)Marshal.ReadInt64((nint)frame);
ulong ret = (ulong)Marshal.ReadInt64((nint)(frame + 8));
if (!TryReadHostQword(frame, out ulong next) || !TryReadHostQword(frame + 8, out ulong ret))
{
Console.Error.WriteLine("[LOADER][WARNING] Could not walk RBP frame chain.");
break;
}
string extra = TryFormatNearestRuntimeSymbol(ret, out string retSym) ? $" [{retSym}]" : string.Empty;
Console.Error.WriteLine($"[LOADER][INFO] frame#{i}: rbp=0x{frame:X16} ret=0x{ret:X16}{extra} next=0x{next:X16}");
if (next <= frame)
@@ -254,10 +338,9 @@ public sealed partial class DirectExecutionBackend
Console.Error.WriteLine("[LOADER][ERROR] - Guest code called an unmapped import");
Console.Error.WriteLine("[LOADER][ERROR] - Guest code accessed unmapped memory");
Console.Error.WriteLine("[LOADER][ERROR] - Need to implement HLE for this NID");
try
byte[] code = new byte[16];
if (TryReadHostBytes(rip, code))
{
byte[] code = new byte[16];
Marshal.Copy((nint)rip, code, 0, code.Length);
Console.Error.WriteLine("[LOADER][INFO] Code at RIP: " + BitConverter.ToString(code).Replace("-", " "));
if (code[0] == 100)
{
@@ -273,25 +356,46 @@ public sealed partial class DirectExecutionBackend
Console.Error.WriteLine($"[LOADER][INFO] RBP: 0x{rbp:X16} (mod 16 = {rbp % 16})");
Console.Error.WriteLine($"[LOADER][INFO] RSP: 0x{rsp:X16} (mod 16 = {rsp % 16})");
}
if (rip > 16)
byte[] before = new byte[16];
if (rip > 16 && TryReadHostBytes(rip - 16, before))
{
byte[] before = new byte[16];
Marshal.Copy((nint)(rip - 16), before, 0, before.Length);
Console.Error.WriteLine("[LOADER][INFO] Code before RIP: " + BitConverter.ToString(before).Replace("-", " "));
}
if (rip > 32)
byte[] window = new byte[64];
if (rip > 32 && TryReadHostBytes(rip - 32, window))
{
byte[] window = new byte[64];
Marshal.Copy((nint)(rip - 32), window, 0, window.Length);
Console.Error.WriteLine("[LOADER][INFO] Code window [RIP-0x20..]: " + BitConverter.ToString(window).Replace("-", " "));
}
for (var stackIndex = 0; stackIndex < 16; stackIndex++)
{
byte[] stackSlot = new byte[8];
if (!TryReadHostBytes(rsp + (ulong)(stackIndex * 8), stackSlot))
{
continue;
}
var candidate = BitConverter.ToUInt64(stackSlot);
if (candidate < _entryPoint || candidate >= _entryPoint + 0x10000000 || candidate < 24)
{
continue;
}
byte[] callSiteWindow = new byte[48];
if (TryReadHostBytes(candidate - 24, callSiteWindow))
{
Console.Error.WriteLine(
$"[LOADER][INFO] Stack guest-code candidate [rsp+0x{stackIndex * 8:X2}]=0x{candidate:X16}, bytes [-0x18..]: " +
BitConverter.ToString(callSiteWindow).Replace("-", " "));
}
}
}
catch
else
{
Console.Error.WriteLine("[LOADER][ERROR] Could not read code at RIP");
}
DumpRecentImportTrace();
DumpGuestDisasmDiagnostics(rip, rbp);
DumpGuestDisasmDiagnostics(rip, rbp, rsp);
DumpGuestRegisterWindowDiagnostics(
rax, rbx, rcx, rdx, rsi, rdi, rbp, rsp,
r8, r9, r10, r11, r12, r13, r14, r15);
DumpGuestReferenceDiagnostics();
DumpGuestPointerWindowDiagnostics();
break;
@@ -299,8 +403,20 @@ public sealed partial class DirectExecutionBackend
Console.Error.WriteLine("[LOADER][WARNING] Type: Breakpoint (int3)");
Console.Error.WriteLine("[LOADER][WARNING] Unexpected breakpoint in direct-bridge mode");
break;
case 1073741845u:
Console.Error.WriteLine("[LOADER][ERROR] Type: Abort (SIGABRT)");
DumpRecentImportTrace();
DumpGuestDisasmDiagnostics(rip, rbp, rsp);
break;
case 3221225501u:
Console.Error.WriteLine("[LOADER][INFO] Type: Illegal Instruction");
byte[] illegalCode = new byte[16];
if (TryReadHostBytes(rip, illegalCode))
{
Console.Error.WriteLine("[LOADER][INFO] Code at RIP: " + BitConverter.ToString(illegalCode).Replace("-", " "));
}
DumpRecentImportTrace();
DumpGuestDisasmDiagnostics(rip, rbp, rsp);
break;
}
@@ -314,6 +430,108 @@ public sealed partial class DirectExecutionBackend
}
}
private unsafe bool TryRecoverAuxiliaryThreadExecuteFault(
EXCEPTION_RECORD* exceptionRecord,
void* contextRecord,
ulong rip)
{
if (exceptionRecord->ExceptionCode != 3221225477u ||
rip >= 0x0000000800000000UL ||
_activeGuestThreadState is not { Name: "tbb_thead" } activeThread)
{
return false;
}
var hostExit = ActiveEntryReturnSentinelRip;
if (hostExit < 0x10000)
{
hostExit = unchecked((ulong)_guestReturnStub);
}
if (hostExit < 0x10000)
{
Console.Error.WriteLine(
$"[LOADER][WARN] Could not recover auxiliary TBB execute fault: target=0x{rip:X16} " +
$"active_exit=0x{ActiveEntryReturnSentinelRip:X16} guest_return_stub=0x{unchecked((ulong)_guestReturnStub):X16}");
return false;
}
_ = TryPatchActiveGuestReturnSlot(hostExit);
WriteCtxU64(contextRecord, 120, 0);
WriteCtxU64(contextRecord, 248, hostExit);
var recovery = Interlocked.Increment(ref _auxiliaryThreadExecuteFaultRecoveries);
Console.Error.WriteLine(
$"[LOADER][WARN] Recovered auxiliary TBB execute fault #{recovery}: " +
$"thread=0x{activeThread.ThreadHandle:X16} target=0x{rip:X16} -> host_exit=0x{hostExit:X16}");
return true;
}
private unsafe bool TryRecoverGuestInt41(uint exceptionCode, void* contextRecord, ulong rip)
{
if (!_ignoreGuestInt41 || exceptionCode != 3221225477u || rip < 0x10000)
{
return false;
}
byte[] opcode = new byte[2];
if (!TryReadHostBytes(rip, opcode) || opcode[0] != 0xCD || opcode[1] != 0x41)
{
return false;
}
var count = Interlocked.Increment(ref _ignoredGuestInt41Count);
WriteCtxU64(contextRecord, 248, rip + 2);
if (count <= 16 || count % 65536 == 0)
{
Console.Error.WriteLine(
$"[LOADER][WARN] Ignored guest int 0x41 trap #{count} at 0x{rip:X16} (default-on; set SHARPEMU_IGNORE_INT41=0 to disable)");
Console.Error.Flush();
}
return true;
}
private unsafe static bool TryRecoverGuestAllocatorHole(
EXCEPTION_RECORD* exceptionRecord,
void* contextRecord,
ulong rip)
{
if (string.Equals(
Environment.GetEnvironmentVariable("SHARPEMU_DISABLE_GUEST_ALLOCATOR_HOLE_RECOVERY"),
"1",
StringComparison.Ordinal) ||
exceptionRecord->NumberParameters < 2 ||
exceptionRecord->ExceptionInformation[0] != 0 ||
exceptionRecord->ExceptionInformation[1] != 8 ||
ReadCtxU64(contextRecord, CTX_RDI) != 0 ||
rip < 0x10000)
{
return false;
}
// Demon's Souls occasionally leaves an empty payload in a locked pool
// tree node. The allocator dereferences payload+8 before reaching its
// existing empty-pool fallback. Match the instruction stream instead of
// a title-specific absolute address, then resume at that fallback so the
// lock is released and the allocator can try its next backing pool.
const ulong allocatorHoleSignature = 0x634CFF568D08778BUL;
if (*(ulong*)rip != allocatorHoleSignature || *((byte*)rip + 8) != 0xF2)
{
return false;
}
const ulong emptyPoolFallbackDelta = 0x8E;
WriteCtxU64(contextRecord, CTX_RIP, rip + emptyPoolFallbackDelta);
var recovery = Interlocked.Increment(ref _guestAllocatorHoleRecoveries);
if (recovery <= 16 || (recovery & (recovery - 1)) == 0)
{
Console.Error.WriteLine(
$"[LOADER][WARN] Guest allocator empty-node adapter recovery #{recovery}: " +
$"rip=0x{rip:X16} -> 0x{rip + emptyPoolFallbackDelta:X16}");
Console.Error.Flush();
}
return true;
}
private static bool IsBenignHostDebugException(uint exceptionCode)
{
return exceptionCode is DBG_PRINTEXCEPTION_C or DBG_PRINTEXCEPTION_WIDE_C or MS_VC_THREADNAME_EXCEPTION;
@@ -393,7 +611,7 @@ public sealed partial class DirectExecutionBackend
}
}
private void DumpGuestDisasmDiagnostics(ulong rip, ulong rbp)
private void DumpGuestDisasmDiagnostics(ulong rip, ulong rbp, ulong rsp)
{
if (!string.Equals(Environment.GetEnvironmentVariable("SHARPEMU_LOG_DISASM"), "1", StringComparison.Ordinal))
{
@@ -405,12 +623,20 @@ public sealed partial class DirectExecutionBackend
DumpGuestInstructionStream("fault-prelude", rip - 0x20, 24);
}
// Optimized guest code frequently omits frame pointers. The return
// address at RSP is then more useful than an RBP walk and identifies the
// exact call site that supplied the faulting arguments.
if (TryReadHostQword(rsp, out var stackReturn) && stackReturn >= 0x60)
{
DumpGuestInstructionStream("stack-return-prelude", stackReturn - 0x60, 40);
}
try
{
ulong frame = rbp;
for (int i = 0; i < 3; i++)
{
if (frame < 140733193388032L || frame > 140737488355327L)
if (frame < 0x10000)
{
break;
}
@@ -554,6 +780,55 @@ public sealed partial class DirectExecutionBackend
}
}
private void DumpGuestRegisterWindowDiagnostics(
ulong rax,
ulong rbx,
ulong rcx,
ulong rdx,
ulong rsi,
ulong rdi,
ulong rbp,
ulong rsp,
ulong r8,
ulong r9,
ulong r10,
ulong r11,
ulong r12,
ulong r13,
ulong r14,
ulong r15)
{
if (!string.Equals(
Environment.GetEnvironmentVariable("SHARPEMU_LOG_REGISTER_WINDOWS"),
"1",
StringComparison.Ordinal))
{
return;
}
// A register can be the only surviving reference to the object or
// argument array that caused a native guest fault. Capture a compact
// window while the process is alive so the post-mortem log can
// distinguish an absent object from a partially initialized one.
var registers = new (string Name, ulong Value)[]
{
("rax", rax), ("rbx", rbx), ("rcx", rcx), ("rdx", rdx),
("rsi", rsi), ("rdi", rdi), ("rbp", rbp), ("rsp", rsp),
("r8", r8), ("r9", r9), ("r10", r10), ("r11", r11),
("r12", r12), ("r13", r13), ("r14", r14), ("r15", r15),
};
var seen = new HashSet<ulong>();
foreach (var (name, value) in registers)
{
if (value < 0x10000 || !seen.Add(value))
{
continue;
}
DumpPointerWindow($"register-{name}", value, 0x80);
}
}
private void ScanExecutableRegionForTargetReferences(
ulong regionBase,
ulong regionEnd,
@@ -821,6 +1096,61 @@ public sealed partial class DirectExecutionBackend
}
}
private static bool TryReadHostQword(ulong address, out ulong value)
{
if (!OperatingSystem.IsWindows())
{
// A stray read inside the signal handler would raise a nested
// SIGSEGV and kill the process before diagnostics finish, so
// probe the region table instead of relying on try/catch.
return TryReadStackU64(address, out value);
}
value = 0;
try
{
value = (ulong)Marshal.ReadInt64((nint)address);
return true;
}
catch
{
return false;
}
}
private unsafe static bool TryReadHostBytes(ulong address, byte[] buffer)
{
if (address < 65536)
{
return false;
}
if (!OperatingSystem.IsWindows())
{
// See TryReadHostQword: probe every touched page before reading.
ulong end = address + (ulong)buffer.Length;
for (ulong page = address & 0xFFFFFFFFFFFFF000uL; page < end; page += 4096)
{
if (VirtualQuery((void*)page, out var mbi, (nuint)sizeof(MEMORY_BASIC_INFORMATION64)) == 0 ||
mbi.State != MEM_COMMIT ||
!IsReadableProtection(mbi.Protect))
{
return false;
}
}
}
try
{
Marshal.Copy((nint)address, buffer, 0, buffer.Length);
return true;
}
catch
{
return false;
}
}
private string FormatPointerWithNearestSymbol(ulong value)
{
string text = $"0x{value:X16}";
@@ -0,0 +1,362 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using System.Buffers.Binary;
using System.Threading;
using Iced.Intel;
using SharpEmu.Core.Cpu.Emulation;
namespace SharpEmu.Core.Cpu.Native;
// Software fallback for the BMI1/BMI2/ABM general-purpose-register instructions.
//
// Guest code runs natively, so the guest and host share the same virtual address space and the
// same registers (the OS delivers them in the CONTEXT record on a fault). When the host CPU lacks
// one of these extensions it raises #UD instead of executing the opcode; without this the title
// simply aborts. Here we decode the faulting instruction, evaluate it against the trapped register
// and memory state, write the result back into the CONTEXT, step RIP past the instruction and ask
// the OS to continue. Only the register-only BMI/ABM forms are handled; anything else returns false
// and falls through to the existing diagnostics unchanged, so this can never mis-handle an opcode it
// does not fully model.
public sealed partial class DirectExecutionBackend
{
// Windows x64 CONTEXT.EFlags lives just past the segment selectors. The GPR offsets it shares
// with the rest of the backend are the CTX_* constants declared in DirectExecutionBackend.cs.
private const int CTX_EFLAGS = 68;
// STATUS_ILLEGAL_INSTRUCTION (#UD surfaced by the Windows vectored handler).
private const uint StatusIllegalInstruction = 0xC000001Du;
private const int MaxInstructionBytes = 15;
// Instruction-window sizes tried in turn so a fault near a page boundary still decodes.
private static readonly int[] DecodeWindowSizes = { MaxInstructionBytes, 11, 8, 4, 2 };
private static int _bmiSoftwareFallbackAnnounced;
private static long _bmiInstructionsEmulated;
private unsafe bool TryRecoverIllegalInstruction(void* contextRecord, ulong rip)
{
if (!TryReadFaultingInstruction(rip, out var instruction))
{
return false;
}
if (instruction.Op0Kind != OpKind.Register ||
!TryGetGprSlot(instruction.Op0Register, out var destOffset, out var size))
{
return false;
}
if (!TryEvaluate(contextRecord, in instruction, size, out var result, out var flagsChanged, out var eflags))
{
return false;
}
WriteCtxU64(contextRecord, destOffset, result);
if (flagsChanged)
{
WriteCtxU32(contextRecord, CTX_EFLAGS, eflags);
}
WriteCtxU64(contextRecord, CTX_RIP, rip + (ulong)instruction.Length);
Interlocked.Increment(ref _bmiInstructionsEmulated);
if (Interlocked.Exchange(ref _bmiSoftwareFallbackAnnounced, 1) == 0)
{
Console.Error.WriteLine(
"[LOADER][INFO] Host lacks a BMI/ABM extension used by the guest; " +
"emulating those instructions in software.");
}
return true;
}
private unsafe bool TryEvaluate(
void* contextRecord,
in Instruction instruction,
GprOperandSize size,
out ulong result,
out bool flagsChanged,
out uint eflags)
{
result = 0;
flagsChanged = false;
eflags = ReadCtxU32(contextRecord, CTX_EFLAGS);
switch (instruction.Mnemonic)
{
case Mnemonic.Andn:
if (!TryReadOperand(contextRecord, in instruction, 1, size, out var andnSrc1) ||
!TryReadOperand(contextRecord, in instruction, 2, size, out var andnSrc2))
{
return false;
}
result = BmiInstructionEmulator.Andn(andnSrc1, andnSrc2, size, ref eflags);
flagsChanged = true;
return true;
case Mnemonic.Blsi:
case Mnemonic.Blsmsk:
case Mnemonic.Blsr:
if (!TryReadOperand(contextRecord, in instruction, 1, size, out var blsSrc))
{
return false;
}
result = instruction.Mnemonic switch
{
Mnemonic.Blsi => BmiInstructionEmulator.Blsi(blsSrc, size, ref eflags),
Mnemonic.Blsmsk => BmiInstructionEmulator.Blsmsk(blsSrc, size, ref eflags),
_ => BmiInstructionEmulator.Blsr(blsSrc, size, ref eflags),
};
flagsChanged = true;
return true;
case Mnemonic.Bextr:
if (!TryReadOperand(contextRecord, in instruction, 1, size, out var bextrSrc) ||
!TryReadOperand(contextRecord, in instruction, 2, size, out var bextrControl))
{
return false;
}
result = BmiInstructionEmulator.Bextr(bextrSrc, bextrControl, size, ref eflags);
flagsChanged = true;
return true;
case Mnemonic.Bzhi:
if (!TryReadOperand(contextRecord, in instruction, 1, size, out var bzhiSrc) ||
!TryReadOperand(contextRecord, in instruction, 2, size, out var bzhiIndex))
{
return false;
}
result = BmiInstructionEmulator.Bzhi(bzhiSrc, bzhiIndex, size, ref eflags);
flagsChanged = true;
return true;
case Mnemonic.Tzcnt:
case Mnemonic.Lzcnt:
if (!TryReadOperand(contextRecord, in instruction, 1, size, out var cntSrc))
{
return false;
}
result = instruction.Mnemonic == Mnemonic.Tzcnt
? BmiInstructionEmulator.Tzcnt(cntSrc, size, ref eflags)
: BmiInstructionEmulator.Lzcnt(cntSrc, size, ref eflags);
flagsChanged = true;
return true;
case Mnemonic.Rorx:
if (instruction.Op2Kind != OpKind.Immediate8 ||
!TryReadOperand(contextRecord, in instruction, 1, size, out var rorxSrc))
{
return false;
}
result = BmiInstructionEmulator.Rorx(rorxSrc, instruction.Immediate8, size);
return true;
case Mnemonic.Sarx:
case Mnemonic.Shlx:
case Mnemonic.Shrx:
if (!TryReadOperand(contextRecord, in instruction, 1, size, out var shiftSrc) ||
!TryReadOperand(contextRecord, in instruction, 2, size, out var shiftCount))
{
return false;
}
result = instruction.Mnemonic switch
{
Mnemonic.Sarx => BmiInstructionEmulator.Sarx(shiftSrc, (int)shiftCount, size),
Mnemonic.Shlx => BmiInstructionEmulator.Shlx(shiftSrc, (int)shiftCount, size),
_ => BmiInstructionEmulator.Shrx(shiftSrc, (int)shiftCount, size),
};
return true;
case Mnemonic.Pdep:
case Mnemonic.Pext:
if (!TryReadOperand(contextRecord, in instruction, 1, size, out var packSrc) ||
!TryReadOperand(contextRecord, in instruction, 2, size, out var packMask))
{
return false;
}
result = instruction.Mnemonic == Mnemonic.Pdep
? BmiInstructionEmulator.Pdep(packSrc, packMask, size)
: BmiInstructionEmulator.Pext(packSrc, packMask, size);
return true;
default:
return false;
}
}
private unsafe bool TryReadFaultingInstruction(ulong rip, out Instruction instruction)
{
// Try the full instruction window first, then shrink so a fault near the end of a mapped
// page (where fewer than 15 bytes are readable) still decodes.
foreach (var attempt in DecodeWindowSizes)
{
var buffer = new byte[attempt];
if (!TryReadHostBytes(rip, buffer))
{
continue;
}
var decoder = Decoder.Create(64, new ByteArrayCodeReader(buffer));
decoder.IP = rip;
decoder.Decode(out instruction);
if (instruction.Code != Code.INVALID && instruction.Length > 0 && instruction.Length <= attempt)
{
return true;
}
}
instruction = default;
return false;
}
private unsafe bool TryReadOperand(
void* contextRecord,
in Instruction instruction,
int operandIndex,
GprOperandSize size,
out ulong value)
{
value = 0;
switch (instruction.GetOpKind(operandIndex))
{
case OpKind.Register:
if (!TryGetGprSlot(instruction.GetOpRegister(operandIndex), out var offset, out _))
{
return false;
}
var raw = ReadCtxU64(contextRecord, offset);
value = size == GprOperandSize.Bits64 ? raw : raw & 0xFFFF_FFFFUL;
return true;
case OpKind.Memory:
if (!TryComputeMemoryAddress(contextRecord, in instruction, out var address))
{
return false;
}
var byteCount = size == GprOperandSize.Bits64 ? 8 : 4;
var buffer = new byte[byteCount];
if (!TryReadHostBytes(address, buffer))
{
return false;
}
value = byteCount == 8
? BinaryPrimitives.ReadUInt64LittleEndian(buffer)
: BinaryPrimitives.ReadUInt32LittleEndian(buffer);
return true;
default:
return false;
}
}
private unsafe bool TryComputeMemoryAddress(void* contextRecord, in Instruction instruction, out ulong address)
{
address = 0;
// FS/GS-relative operands need the guest segment base, which is not modelled here.
if (instruction.SegmentPrefix != Register.None)
{
return false;
}
if (instruction.IsIPRelativeMemoryOperand)
{
address = instruction.IPRelativeMemoryAddress;
return true;
}
var effective = instruction.MemoryDisplacement64;
if (instruction.MemoryBase != Register.None)
{
if (!TryGetGpr64Offset(instruction.MemoryBase, out var baseOffset))
{
return false;
}
effective += ReadCtxU64(contextRecord, baseOffset);
}
if (instruction.MemoryIndex != Register.None)
{
if (!TryGetGpr64Offset(instruction.MemoryIndex, out var indexOffset))
{
return false;
}
effective += ReadCtxU64(contextRecord, indexOffset) * (ulong)instruction.MemoryIndexScale;
}
address = effective;
return true;
}
// Maps a 32- or 64-bit GPR to its CONTEXT offset and reports the operand width it implies.
private static bool TryGetGprSlot(Register register, out int offset, out GprOperandSize size)
{
switch (register)
{
case Register.EAX: offset = CTX_RAX; size = GprOperandSize.Bits32; return true;
case Register.ECX: offset = CTX_RCX; size = GprOperandSize.Bits32; return true;
case Register.EDX: offset = CTX_RDX; size = GprOperandSize.Bits32; return true;
case Register.EBX: offset = CTX_RBX; size = GprOperandSize.Bits32; return true;
case Register.ESP: offset = CTX_RSP; size = GprOperandSize.Bits32; return true;
case Register.EBP: offset = CTX_RBP; size = GprOperandSize.Bits32; return true;
case Register.ESI: offset = CTX_RSI; size = GprOperandSize.Bits32; return true;
case Register.EDI: offset = CTX_RDI; size = GprOperandSize.Bits32; return true;
case Register.R8D: offset = CTX_R8; size = GprOperandSize.Bits32; return true;
case Register.R9D: offset = CTX_R9; size = GprOperandSize.Bits32; return true;
case Register.R10D: offset = CTX_R10; size = GprOperandSize.Bits32; return true;
case Register.R11D: offset = CTX_R11; size = GprOperandSize.Bits32; return true;
case Register.R12D: offset = CTX_R12; size = GprOperandSize.Bits32; return true;
case Register.R13D: offset = CTX_R13; size = GprOperandSize.Bits32; return true;
case Register.R14D: offset = CTX_R14; size = GprOperandSize.Bits32; return true;
case Register.R15D: offset = CTX_R15; size = GprOperandSize.Bits32; return true;
default:
if (TryGetGpr64Offset(register, out offset))
{
size = GprOperandSize.Bits64;
return true;
}
size = GprOperandSize.Bits32;
return false;
}
}
private static bool TryGetGpr64Offset(Register register, out int offset)
{
switch (register)
{
case Register.RAX: offset = CTX_RAX; return true;
case Register.RCX: offset = CTX_RCX; return true;
case Register.RDX: offset = CTX_RDX; return true;
case Register.RBX: offset = CTX_RBX; return true;
case Register.RSP: offset = CTX_RSP; return true;
case Register.RBP: offset = CTX_RBP; return true;
case Register.RSI: offset = CTX_RSI; return true;
case Register.RDI: offset = CTX_RDI; return true;
case Register.R8: offset = CTX_R8; return true;
case Register.R9: offset = CTX_R9; return true;
case Register.R10: offset = CTX_R10; return true;
case Register.R11: offset = CTX_R11; return true;
case Register.R12: offset = CTX_R12; return true;
case Register.R13: offset = CTX_R13; return true;
case Register.R14: offset = CTX_R14; return true;
case Register.R15: offset = CTX_R15; return true;
default: offset = 0; return false;
}
}
}
File diff suppressed because it is too large Load Diff
@@ -92,7 +92,10 @@ public sealed partial class DirectExecutionBackend
private NativeGuestExecutor? RentNativeGuestExecutor()
{
if (NativeGuestWorkersDisabled)
// NativeGuestExecutor emits a Win32 wait loop and creates it with
// kernel32!CreateThread. POSIX hosts use the established inline entry
// path until the worker loop has a pthread/eventfd implementation.
if (!OperatingSystem.IsWindows() || NativeGuestWorkersDisabled)
{
return null;
}
@@ -185,8 +188,20 @@ public sealed partial class DirectExecutionBackend
private static nint _exitThreadAddress;
private readonly DirectExecutionBackend _backend;
private readonly AutoResetEvent _workAvailable = new(false);
private readonly AutoResetEvent _workCompleted = new(false);
// Windows uses AutoResetEvent (its SafeWaitHandle is a real kernel
// event the emitted loop can wait on); POSIX uses worker-event
// semaphores shared the same way via PosixHostStubs.
private readonly AutoResetEvent? _workAvailable;
private readonly AutoResetEvent? _workCompleted;
private nint _workSemaphore;
private nint _doneSemaphore;
// RunPrologue/RunEpilogue compile to the host ABI (SysV on POSIX); the
// emitted loop calls them with Win64 registers, so POSIX routes the
// calls through register-shuffling thunks (shared by all workers).
private static nint _posixPrologueThunk;
private static nint _posixEpilogueThunk;
private static readonly object PosixThunkGate = new();
private GCHandle _selfHandle;
private void* _controlBlock;
private void* _loopStub;
@@ -225,6 +240,11 @@ public sealed partial class DirectExecutionBackend
private NativeGuestExecutor(DirectExecutionBackend backend)
{
_backend = backend;
if (OperatingSystem.IsWindows())
{
_workAvailable = new AutoResetEvent(false);
_workCompleted = new AutoResetEvent(false);
}
}
public static NativeGuestExecutor? TryCreate(DirectExecutionBackend backend)
@@ -276,8 +296,34 @@ public sealed partial class DirectExecutionBackend
var prologuePtr = (nint)(delegate* unmanaged<nint, nint>)&RunPrologue;
var epiloguePtr = (nint)(delegate* unmanaged<nint, int, void>)&RunEpilogue;
var executorHandle = GCHandle.ToIntPtr(_selfHandle);
var workHandle = _workAvailable.SafeWaitHandle.DangerousGetHandle();
var doneHandle = _workCompleted.SafeWaitHandle.DangerousGetHandle();
nint workHandle;
nint doneHandle;
if (OperatingSystem.IsWindows())
{
workHandle = _workAvailable!.SafeWaitHandle.DangerousGetHandle();
doneHandle = _workCompleted!.SafeWaitHandle.DangerousGetHandle();
}
else
{
lock (PosixThunkGate)
{
if (_posixPrologueThunk == 0)
{
_posixPrologueThunk = PosixHostStubs.CreateWin64ToSysVThunk(prologuePtr);
_posixEpilogueThunk = PosixHostStubs.CreateWin64ToSysVThunk(epiloguePtr);
}
}
prologuePtr = _posixPrologueThunk;
epiloguePtr = _posixEpilogueThunk;
_workSemaphore = PosixHostStubs.CreateWorkerEvent();
_doneSemaphore = PosixHostStubs.CreateWorkerEvent();
if (_workSemaphore == 0 || _doneSemaphore == 0)
{
return false;
}
workHandle = _workSemaphore;
doneHandle = _doneSemaphore;
}
byte* code = (byte*)_loopStub;
int offset = 0;
@@ -397,8 +443,8 @@ public sealed partial class DirectExecutionBackend
_runYieldRequested = false;
_runYieldReason = null;
_runForcedExit = false;
_workAvailable.Set();
_workCompleted.WaitOne();
SignalWorkAvailable();
WaitWorkCompleted();
_runContext = null;
_runState = null;
yieldRequested = _runYieldRequested;
@@ -411,6 +457,28 @@ public sealed partial class DirectExecutionBackend
return _runNativeResult;
}
private void SignalWorkAvailable()
{
if (_workAvailable is not null)
{
_workAvailable.Set();
return;
}
_ = PosixHostStubs.SignalWorkerEvent(_workSemaphore);
}
private void WaitWorkCompleted()
{
if (_workCompleted is not null)
{
_workCompleted.WaitOne();
return;
}
_ = PosixHostStubs.WaitWorkerEvent(_doneSemaphore, -1);
}
[UnmanagedCallersOnly]
private static nint RunPrologue(nint executorHandle)
{
@@ -540,7 +608,7 @@ public sealed partial class DirectExecutionBackend
}
try
{
_workAvailable.Set();
SignalWorkAvailable();
}
catch (ObjectDisposedException)
{
@@ -575,8 +643,18 @@ public sealed partial class DirectExecutionBackend
{
_selfHandle.Free();
}
_workAvailable.Dispose();
_workCompleted.Dispose();
_workAvailable?.Dispose();
_workCompleted?.Dispose();
if (_workSemaphore != 0)
{
PosixHostStubs.DestroyWorkerEvent(_workSemaphore);
_workSemaphore = 0;
}
if (_doneSemaphore != 0)
{
PosixHostStubs.DestroyWorkerEvent(_doneSemaphore);
_doneSemaphore = 0;
}
}
}
}
@@ -0,0 +1,459 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using System;
using System.Runtime.InteropServices;
using System.Threading;
using SharpEmu.HLE;
namespace SharpEmu.Core.Cpu.Native;
public sealed unsafe partial class DirectExecutionBackend
{
// POSIX bridge for the Windows vectored-exception-handler logic. A
// sigaction(SIGSEGV/SIGBUS/SIGILL) handler rebuilds the EXCEPTION_POINTERS
// view the shared handlers expect (Win64 CONTEXT register offsets) from
// the signal's mcontext, runs the same recovery chain the VEH path uses
// (unresolved-import trap sentinels, demand-paging of lazily-committed
// guest pages, fault diagnostics), and writes register changes back into
// the mcontext so sigreturn resumes the repaired guest. Unrecovered
// faults are forwarded to the previously installed handler so the .NET
// runtime keeps turning its own faults into managed exceptions.
private const int PosixSigIll = 4;
private const int PosixSigTrap = 5;
private const int PosixSigAbort = 6;
private const int PosixSigSegv = 11;
private static readonly int PosixSigBus = OperatingSystem.IsMacOS() ? 10 : 7;
// struct sigaction: the handler pointer leads on both platforms; Darwin
// packs { handler(8), mask(4), flags(4) }, Linux glibc/musl packs
// { handler(8), mask(128), flags(4), restorer(8) }.
private static readonly int PosixSigactionSize = OperatingSystem.IsMacOS() ? 16 : 152;
private static readonly int PosixSigactionFlagsOffset = OperatingSystem.IsMacOS() ? 12 : 136;
private static readonly int PosixSaSigInfo = OperatingSystem.IsMacOS() ? 0x0040 : 0x0004;
private static readonly int PosixSaNoDefer = OperatingSystem.IsMacOS() ? 0x0010 : 0x40000000;
// siginfo_t.si_addr: Darwin { signo, errno, code, pid, uid, status, addr },
// Linux { signo, errno, code, pad32, addr }.
private static readonly int PosixSigInfoAddressOffset = OperatingSystem.IsMacOS() ? 24 : 16;
// Darwin ucontext_t stores a pointer to __darwin_mcontext64 at +48; the
// general registers live in its __ss thread state after the 16-byte
// exception state. Linux glibc embeds mcontext_t inline at +40 with the
// registers in gregs[23]. Rosetta 2 delivers the regular x86-64 layout
// to translated processes.
private const int DarwinUcontextMcontextOffset = 48;
private const int DarwinMcontextErrOffset = 4;
private const int DarwinMcontextFaultAddressOffset = 8;
private const int LinuxUcontextGregsOffset = 40;
private const int LinuxGregsErrOffset = 19 * 8;
// The kernel's x86-64 sigcontext places the FXSAVE-image pointer right
// after the general registers it hands to the handler: err(152)
// trapno(160) oldmask(168) cr2(176) fpstate(184), all relative to
// GetPosixRegisterBase. glibc and musl both overlay this kernel layout
// verbatim (glibc's mcontext_t.fpregs is the same slot), so the offset
// is libc-independent. Inside the FXSAVE image the XMM registers start
// at +160 (32-byte header + 8 legacy x87/MMX slots x 16 bytes) - the
// same relative position they occupy in the Win64 CONTEXT's FltSave
// area (Win64ContextXmm0Offset = 256 + 160).
private const int LinuxGregsFpstateOffset = 184;
private const int FxsaveXmmOffset = 160;
private const int XmmBlockSize = 16 * 16;
// Byte offsets of the general registers relative to GetPosixRegisterBase,
// ordered to match the contiguous Win64 CONTEXT block CTX_RAX..CTX_RIP
// (rax, rcx, rdx, rbx, rsp, rbp, rsi, rdi, r8..r15, rip). Verified
// against the x86-64 platform headers.
private static readonly int[] PosixRegisterOffsets = OperatingSystem.IsMacOS()
? new[] { 16, 32, 40, 24, 72, 64, 56, 48, 80, 88, 96, 104, 112, 120, 128, 136, 144 }
: new[] { 104, 112, 96, 88, 120, 80, 72, 64, 0, 8, 16, 24, 32, 40, 48, 56, 128 };
private static DirectExecutionBackend? _posixSignalBackend;
private static bool _posixSignalHandlersInstalled;
private static bool _posixRawRecoveryEnabled;
private static bool _posixSignalWarmup;
private static readonly nint[] _posixPreviousActions = new nint[32];
private static int _posixSignalTraceCount;
private static long _perfSignalCount;
private static readonly bool _perfSignalCounter =
string.Equals(Environment.GetEnvironmentVariable("SHARPEMU_PERF_MEM"), "1", StringComparison.Ordinal);
[ThreadStatic]
private static int _posixSignalHandlerDepth;
// True while the current thread's in-flight POSIX fault carries the real
// XMM registers in the CONTEXT scratch buffer and writes to them will
// reach the mcontext on resume. Gates recovery paths (SSE4a EXTRQ/
// INSERTQ) that would otherwise compute results from a zeroed XMM area
// and silently discard what they "wrote". Darwin is not bridged yet, so
// the flag stays false there.
[ThreadStatic]
private static bool _posixXmmContextBridged;
private void SetupPosixExceptionHandler()
{
if (string.Equals(Environment.GetEnvironmentVariable("SHARPEMU_DISABLE_POSIX_SIGNALS"), "1", StringComparison.Ordinal))
{
Console.Error.WriteLine("[LOADER][WARN] POSIX signal exception bridge disabled by SHARPEMU_DISABLE_POSIX_SIGNALS=1; guest faults will not be recovered.");
return;
}
_posixSignalBackend = this;
if (_posixSignalHandlersInstalled)
{
return;
}
_posixRawRecoveryEnabled = !string.Equals(
Environment.GetEnvironmentVariable("SHARPEMU_DISABLE_RAW_HANDLER"), "1", StringComparison.Ordinal);
if (!_posixRawRecoveryEnabled)
{
Console.Error.WriteLine("[LOADER][INFO] Raw sentinel recovery disabled by SHARPEMU_DISABLE_RAW_HANDLER=1");
}
WarmUpPosixSignalPath();
SharpEmu.HLE.GuestImageWriteTracker.WarmUp();
if (!InstallPosixSignalHandler(PosixSigSegv) ||
!InstallPosixSignalHandler(PosixSigBus) ||
!InstallPosixSignalHandler(PosixSigIll) ||
!InstallPosixSignalHandler(PosixSigTrap) ||
!InstallPosixSignalHandler(PosixSigAbort))
{
throw new InvalidOperationException("Failed to install POSIX fault signal handlers");
}
_posixSignalHandlersInstalled = true;
Console.Error.WriteLine("[LOADER][INFO] POSIX signal exception bridge installed (SIGSEGV/SIGBUS/SIGILL)");
}
/// <summary>
/// Runs the signal-recovery path once with fabricated inputs before the
/// handlers are installed. The first entry into the handler must not
/// require JIT compilation (a fault can interrupt arbitrary runtime
/// states), and under Rosetta 2 the signal trampoline cannot enter x86
/// code that has never been executed (and therefore never translated): a
/// cold handler is silently never invoked and the faulting instruction
/// retries forever.
/// </summary>
private void WarmUpPosixSignalPath()
{
byte* fakeUcontext = stackalloc byte[512];
new Span<byte>(fakeUcontext, 512).Clear();
byte* fakeMcontext = stackalloc byte[512];
new Span<byte>(fakeMcontext, 512).Clear();
if (OperatingSystem.IsMacOS())
{
*(byte**)(fakeUcontext + DarwinUcontextMcontextOffset) = fakeMcontext;
}
_posixSignalWarmup = true;
try
{
((delegate* unmanaged<int, nint, nint, void>)&HandlePosixSignal)(PosixSigSegv, 0, (nint)fakeUcontext);
// Warm the branches the fabricated fault above skips without
// spamming diagnostics: the benign-exception path through
// VectoredHandler, the lazy-commit probe (fault address 0 bails
// out immediately), and the chain helper (signal 0 has no saved
// action and sigaction(0, ...) fails with EINVAL).
EXCEPTION_RECORD record = default;
record.ExceptionCode = DBG_PRINTEXCEPTION_C;
byte* contextRecord = stackalloc byte[Win64ContextSize];
new Span<byte>(contextRecord, Win64ContextSize).Clear();
EXCEPTION_POINTERS pointers;
pointers.ExceptionRecord = &record;
pointers.ContextRecord = contextRecord;
_ = VectoredHandler(&pointers);
record.ExceptionCode = 3221225477u;
record.NumberParameters = 2;
// 0x70000 is never guest-owned, so this walks the vmem region
// scan and the PRT range check, then bails out silently.
record.ExceptionInformation[1] = 0x70000;
_ = TryHandleLazyCommittedPage(&record, 0, 0);
ChainPreviousPosixAction(0, 0, 0);
}
finally
{
_posixSignalWarmup = false;
}
}
private static bool InstallPosixSignalHandler(int signal)
{
byte* action = stackalloc byte[PosixSigactionSize];
new Span<byte>(action, PosixSigactionSize).Clear();
*(nint*)action = (nint)(delegate* unmanaged<int, nint, nint, void>)&HandlePosixSignal;
// No SA_ONSTACK: the runtime's alternate stacks are far too small for
// the recovery/diagnostic path (JIT compilation of cold handler code
// can run inside the signal frame). Guest faults deliver onto the 2MB
// guest stack, host faults onto the regular thread stack — the same
// stacks Windows dispatches exceptions on.
*(int*)(action + PosixSigactionFlagsOffset) = PosixSaSigInfo | PosixSaNoDefer;
var previous = (byte*)NativeMemory.AllocZeroed((nuint)PosixSigactionSize);
if (sigaction(signal, action, previous) != 0)
{
NativeMemory.Free(previous);
Console.Error.WriteLine($"[LOADER][ERROR] sigaction({signal}) failed: errno={Marshal.GetLastPInvokeError()}");
return false;
}
_posixPreviousActions[signal] = (nint)previous;
return true;
}
[UnmanagedCallersOnly]
private static void HandlePosixSignal(int signal, nint siginfo, nint ucontext)
{
if (_posixSignalHandlerDepth > 0)
{
// A fault inside our own fault handler (diagnostics touched an
// unmapped address): restore the default action and return so the
// re-executed instruction terminates the process.
RestoreDefaultPosixAction(signal);
return;
}
_posixSignalHandlerDepth++;
if (_perfSignalCounter)
{
var n = Interlocked.Increment(ref _perfSignalCount);
if (n % 100000 == 0)
{
Console.Error.WriteLine($"[PERF][MEM] posix_faults={n}");
}
}
try
{
// Guest-image write tracking runs first: it only needs the fault
// address (safe for host and guest threads alike) and must resume
// the faulting write immediately after restoring write access.
if (signal != PosixSigIll &&
siginfo != 0 &&
SharpEmu.HLE.GuestImageWriteTracker.TryHandleWriteFault(
*(ulong*)((byte*)siginfo + PosixSigInfoAddressOffset)))
{
return;
}
if (TryHandlePosixFault(signal, siginfo, ucontext))
{
return;
}
}
catch
{
// A managed exception must never unwind out of a signal frame.
}
finally
{
_posixSignalHandlerDepth--;
}
ChainPreviousPosixAction(signal, siginfo, ucontext);
}
private static bool TryHandlePosixFault(int signal, nint siginfo, nint ucontext)
{
byte* registers = GetPosixRegisterBase(ucontext);
if (registers == null)
{
return false;
}
byte* contextRecord = stackalloc byte[Win64ContextSize];
new Span<byte>(contextRecord, Win64ContextSize).Clear();
int[] offsets = PosixRegisterOffsets;
for (int i = 0; i < offsets.Length; i++)
{
WriteCtxU64(contextRecord, CTX_RAX + i * 8, *(ulong*)(registers + offsets[i]));
}
// Bridge the XMM registers alongside the GPRs where the layout is
// known: on Linux the fpstate pointer and FXSAVE image are kernel
// ABI, so recovery paths that read or write XMM state (SSE4a
// EXTRQ/INSERTQ) see the live registers and their writes reach the
// guest through sigreturn.
byte* fpstate = null;
if (OperatingSystem.IsLinux())
{
fpstate = *(byte**)(registers + LinuxGregsFpstateOffset);
if (fpstate != null)
{
Buffer.MemoryCopy(
fpstate + FxsaveXmmOffset,
contextRecord + Win64ContextXmm0Offset,
XmmBlockSize,
XmmBlockSize);
}
}
_posixXmmContextBridged = fpstate != null;
EXCEPTION_RECORD record = default;
record.ExceptionAddress = (void*)ReadCtxU64(contextRecord, CTX_RIP);
if (signal == PosixSigIll)
{
record.ExceptionCode = 3221225501u;
}
else if (signal == PosixSigTrap)
{
record.ExceptionCode = 2147483651u;
}
else if (signal == PosixSigAbort)
{
record.ExceptionCode = 1073741845u;
}
else
{
ulong faultAddress = GetPosixFaultAddress(siginfo, registers);
record.ExceptionCode = 3221225477u;
record.NumberParameters = 2;
record.ExceptionInformation[0] = GetPosixAccessType(registers, faultAddress, ReadCtxU64(contextRecord, CTX_RIP));
record.ExceptionInformation[1] = faultAddress;
}
EXCEPTION_POINTERS pointers;
pointers.ExceptionRecord = &record;
pointers.ContextRecord = contextRecord;
int traceIndex = _posixSignalWarmup ? 0 : Interlocked.Increment(ref _posixSignalTraceCount);
bool traceSignal = traceIndex > 0 && (traceIndex <= 16 || traceIndex % 1024 == 0 ||
string.Equals(Environment.GetEnvironmentVariable("SHARPEMU_LOG_POSIX_SIGNALS"), "1", StringComparison.Ordinal));
if (traceSignal)
{
Console.Error.WriteLine(
$"[LOADER][TRACE] posix-signal#{traceIndex}: sig={signal} rip=0x{ReadCtxU64(contextRecord, CTX_RIP):X16} " +
$"fault=0x{record.ExceptionInformation[1]:X16} access={record.ExceptionInformation[0]} rsp=0x{ReadCtxU64(contextRecord, CTX_RSP):X16}");
Console.Error.Flush();
}
// Sentinel recovery runs first: on Windows both vectored handlers see
// every fault anyway, and recovering here avoids dumping the full
// VectoredHandler diagnostics for each recoverable trap.
int disposition = 0;
if (_posixRawRecoveryEnabled)
{
disposition = TryRecoverUnresolvedSentinel(&pointers);
}
if (disposition != -1 && !_posixSignalWarmup && _posixSignalBackend is { } backend)
{
disposition = backend.VectoredHandler(&pointers);
}
if (traceSignal)
{
Console.Error.WriteLine(
$"[LOADER][TRACE] posix-signal#{traceIndex}: recovered={disposition == -1} new_rip=0x{ReadCtxU64(contextRecord, CTX_RIP):X16}");
Console.Error.Flush();
}
if (disposition != -1 && !_posixSignalWarmup)
{
return false;
}
for (int i = 0; i < offsets.Length; i++)
{
*(ulong*)(registers + offsets[i]) = ReadCtxU64(contextRecord, CTX_RAX + i * 8);
}
if (fpstate != null)
{
Buffer.MemoryCopy(
contextRecord + Win64ContextXmm0Offset,
fpstate + FxsaveXmmOffset,
XmmBlockSize,
XmmBlockSize);
}
return true;
}
private static byte* GetPosixRegisterBase(nint ucontext)
{
if (ucontext == 0)
{
return null;
}
if (OperatingSystem.IsMacOS())
{
return *(byte**)((byte*)ucontext + DarwinUcontextMcontextOffset);
}
return (byte*)ucontext + LinuxUcontextGregsOffset;
}
private static ulong GetPosixFaultAddress(nint siginfo, byte* registers)
{
ulong address = siginfo != 0 ? *(ulong*)((byte*)siginfo + PosixSigInfoAddressOffset) : 0;
if (address == 0 && OperatingSystem.IsMacOS())
{
address = *(ulong*)(registers + DarwinMcontextFaultAddressOffset);
}
return address;
}
private static ulong GetPosixAccessType(byte* registers, ulong faultAddress, ulong rip)
{
// x86 page-fault error code: bit 1 = write access, bit 4 = instruction
// fetch. Fall back to comparing the fault address against RIP when
// the error code is not populated (e.g. under Rosetta 2 translation).
ulong error = OperatingSystem.IsMacOS()
? *(uint*)(registers + DarwinMcontextErrOffset)
: *(ulong*)(registers + LinuxGregsErrOffset);
if ((error & 0x10) != 0)
{
return 8;
}
if ((error & 0x2) != 0)
{
return 1;
}
return faultAddress != 0 && faultAddress == rip ? 8u : 0u;
}
private static void RestoreDefaultPosixAction(int signal)
{
byte* action = stackalloc byte[PosixSigactionSize];
new Span<byte>(action, PosixSigactionSize).Clear();
_ = sigaction(signal, action, null);
}
private static void ChainPreviousPosixAction(int signal, nint siginfo, nint ucontext)
{
byte* previous = (uint)signal < (uint)_posixPreviousActions.Length
? (byte*)_posixPreviousActions[signal]
: null;
nint handler = previous != null ? *(nint*)previous : 0;
if (handler == 0)
{
// SIG_DFL (or nothing saved): reinstate the default action and
// return, so re-executing the faulting instruction terminates the
// process with the original fault context intact.
RestoreDefaultPosixAction(signal);
return;
}
if (handler == 1)
{
// SIG_IGN
return;
}
int flags = *(int*)(previous + PosixSigactionFlagsOffset);
if ((flags & PosixSaSigInfo) != 0)
{
((delegate* unmanaged<int, nint, nint, void>)handler)(signal, siginfo, ucontext);
}
else
{
((delegate* unmanaged<int, void>)handler)(signal);
}
}
[DllImport("libc", SetLastError = true)]
private static extern int sigaction(int signum, void* act, void* oldact);
}
File diff suppressed because it is too large Load Diff
+1 -1
View File
@@ -252,7 +252,7 @@ public static unsafe class JitStubs
var pattern = TlsAccessPattern;
var end = start + length - pattern.Length;
for (var ptr = start; ptr < end; ptr++)
for (var ptr = start; ptr <= end; ptr++)
{
if (MatchesPattern(ptr, pattern))
{
@@ -0,0 +1,49 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using SharpEmu.HLE.Host;
namespace SharpEmu.Core.Cpu.Native;
/// <summary>
/// Placeholder for hosts whose fault bridge is installed directly by the
/// execution backend. POSIX uses its sigaction bridge and never calls these
/// Windows-shaped registration methods.
/// </summary>
internal sealed class NullHostFaultHandling : IHostFaultHandling
{
public static NullHostFaultHandling Instance { get; } = new();
private NullHostFaultHandling()
{
}
public nint CreateHandlerThunk(nint managedCallback, uint hostRspSwitchTlsSlot, nint tlsGetValueAddress)
{
_ = managedCallback;
_ = hostRspSwitchTlsSlot;
_ = tlsGetValueAddress;
return 0;
}
public void FreeThunk(nint thunk)
{
_ = thunk;
}
public nint AddFirstChanceHandler(nint thunk)
{
_ = thunk;
return 0;
}
public void RemoveHandler(nint handle)
{
_ = handle;
}
public void SetUnhandledFilter(nint thunk)
{
_ = thunk;
}
}
@@ -0,0 +1,425 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using System.Runtime.InteropServices;
using SharpEmu.HLE;
namespace SharpEmu.Core.Cpu.Native;
/// <summary>
/// POSIX replacements for the kernel32 helpers the native backend embeds in
/// emitted x86-64 code. Every stub exposed here follows the Win64 calling
/// convention the emitted call sites were written for (first argument in
/// ECX, result in RAX, Win64 non-volatile registers preserved), so the
/// emission code stays identical across platforms.
/// </summary>
internal static unsafe class PosixHostStubs
{
private static readonly object Gate = new();
private static bool _initialized;
private static nint _tlsGetValueStub;
private static nint _queryPerformanceCounterStub;
private static nint _switchToThreadStub;
private static nint _sleepStub;
public static nint TlsGetValueStubAddress
{
get { EnsureInitialized(); return _tlsGetValueStub; }
}
public static nint QueryPerformanceCounterStubAddress
{
get { EnsureInitialized(); return _queryPerformanceCounterStub; }
}
public static nint SwitchToThreadStubAddress
{
get { EnsureInitialized(); return _switchToThreadStub; }
}
public static nint SleepStubAddress
{
get { EnsureInitialized(); return _sleepStub; }
}
public static nint CreateWorkerEvent()
{
if (OperatingSystem.IsMacOS())
{
return dispatch_semaphore_create(0);
}
var semaphore = Marshal.AllocHGlobal(64);
if (sem_init(semaphore, 0, 0) != 0)
{
Marshal.FreeHGlobal(semaphore);
return 0;
}
return semaphore;
}
public static bool SignalWorkerEvent(nint handle)
{
if (OperatingSystem.IsMacOS())
{
_ = dispatch_semaphore_signal(handle);
return true;
}
return sem_post(handle) == 0;
}
public static bool WaitWorkerEvent(nint handle, int timeoutMilliseconds)
{
if (OperatingSystem.IsMacOS())
{
if (timeoutMilliseconds < 0)
{
return dispatch_semaphore_wait(handle, ulong.MaxValue) == 0;
}
var deadline = dispatch_time(0, timeoutMilliseconds * 1_000_000L);
return dispatch_semaphore_wait(handle, deadline) == 0;
}
if (timeoutMilliseconds < 0)
{
while (sem_wait(handle) != 0)
{
}
return true;
}
var deadlineTicks = Environment.TickCount64 + timeoutMilliseconds;
while (sem_trywait(handle) != 0)
{
if (Environment.TickCount64 >= deadlineTicks)
{
return false;
}
Thread.Sleep(1);
}
return true;
}
public static void DestroyWorkerEvent(nint handle)
{
if (handle == 0)
{
return;
}
if (OperatingSystem.IsMacOS())
{
dispatch_release(handle);
return;
}
_ = sem_destroy(handle);
Marshal.FreeHGlobal(handle);
}
/// <summary>Allocates a pthread TLS key, mirroring kernel32!TlsAlloc.</summary>
public static uint TlsAlloc()
{
if (OperatingSystem.IsMacOS())
{
nuint key;
return pthread_key_create_mac(&key, 0) == 0 ? (uint)key : uint.MaxValue;
}
uint key32;
return pthread_key_create_linux(&key32, 0) == 0 ? key32 : uint.MaxValue;
}
public static bool TlsFree(uint key)
{
return OperatingSystem.IsMacOS()
? pthread_key_delete_mac((nuint)key) == 0
: pthread_key_delete_linux(key) == 0;
}
public static bool TlsSetValue(uint key, nint value)
{
return OperatingSystem.IsMacOS()
? pthread_setspecific_mac((nuint)key, value) == 0
: pthread_setspecific_linux(key, value) == 0;
}
public static nint TlsGetValue(uint key)
{
return OperatingSystem.IsMacOS()
? pthread_getspecific_mac((nuint)key)
: pthread_getspecific_linux(key);
}
/// <summary>Stable numeric id of the calling thread (kernel32!GetCurrentThreadId).</summary>
public static uint GetCurrentThreadId()
{
if (OperatingSystem.IsMacOS())
{
ulong tid;
return pthread_threadid_np(0, &tid) == 0 ? unchecked((uint)tid) : 0u;
}
return unchecked((uint)gettid());
}
/// <summary>
/// Wraps a managed callback (compiled for the SysV ABI on POSIX .NET) in a
/// thunk that accepts up to four integer arguments in the Win64 ABI the
/// emitted x86-64 call sites use. Win64 passes args in rcx/rdx/r8/r9 and
/// treats rdi/rsi as non-volatile; SysV expects rdi/rsi/rdx/rcx and
/// clobbers them, so the thunk saves rdi/rsi, shuffles the registers, keeps
/// the stack 16-byte aligned for the call, and forwards the rax result.
/// </summary>
public static nint CreateWin64ToSysVThunk(nint sysvTarget)
{
var page = (byte*)HostMemory.Alloc(
null,
4096,
HostMemory.MEM_COMMIT | HostMemory.MEM_RESERVE,
HostMemory.PAGE_EXECUTE_READWRITE);
if (page == null)
{
throw new OutOfMemoryException("Failed to allocate Win64->SysV thunk page");
}
var offset = 0;
Emit(page, ref offset, 0x57); // push rdi
Emit(page, ref offset, 0x56); // push rsi
Emit(page, ref offset, 0x48, 0x89, 0xCF); // mov rdi, rcx
Emit(page, ref offset, 0x48, 0x89, 0xD6); // mov rsi, rdx
Emit(page, ref offset, 0x4C, 0x89, 0xC2); // mov rdx, r8
Emit(page, ref offset, 0x4C, 0x89, 0xC9); // mov rcx, r9
Emit(page, ref offset, 0x48, 0x83, 0xEC, 0x08); // sub rsp, 8 (realign to 16)
EmitMovRaxImm64(page, ref offset, sysvTarget); // mov rax, target
Emit(page, ref offset, 0xFF, 0xD0); // call rax
Emit(page, ref offset, 0x48, 0x83, 0xC4, 0x08); // add rsp, 8
Emit(page, ref offset, 0x5E); // pop rsi
Emit(page, ref offset, 0x5F); // pop rdi
Emit(page, ref offset, 0xC3); // ret
if (!HostMemory.Protect(page, 4096, HostMemory.PAGE_EXECUTE_READ, out _))
{
throw new InvalidOperationException("Failed to protect Win64->SysV thunk page");
}
HostMemory.FlushInstructionCache(page, (nuint)offset);
return (nint)page;
}
private static void EnsureInitialized()
{
if (_initialized)
{
return;
}
lock (Gate)
{
if (_initialized)
{
return;
}
BuildStubs();
_initialized = true;
}
}
private static void BuildStubs()
{
var page = (byte*)HostMemory.Alloc(
null,
4096,
HostMemory.MEM_COMMIT | HostMemory.MEM_RESERVE,
HostMemory.PAGE_EXECUTE_READWRITE);
if (page == null)
{
throw new OutOfMemoryException("Failed to allocate POSIX host helper stub page");
}
var offset = 0;
_tlsGetValueStub = EmitTlsGetValue(page, ref offset);
_queryPerformanceCounterStub = EmitQueryPerformanceCounter(page, ref offset);
_switchToThreadStub = EmitSwitchToThread(page, ref offset);
_sleepStub = EmitSleep(page, ref offset);
if (!HostMemory.Protect(page, 4096, HostMemory.PAGE_EXECUTE_READ, out _))
{
throw new InvalidOperationException("Failed to protect POSIX host helper stub page");
}
HostMemory.FlushInstructionCache(page, (nuint)offset);
}
private static nint EmitTlsGetValue(byte* page, ref int offset)
{
var start = (nint)(page + offset);
if (OperatingSystem.IsMacOS())
{
// On macOS x86-64 pthread keys index the gs-based thread specific
// data array directly, so TlsGetValue(index in ecx) collapses to a
// single load that clobbers nothing but RAX.
Emit(page, ref offset, 0x89, 0xC8); // mov eax, ecx
Emit(page, ref offset, 0x65, 0x48, 0x8B, 0x04, 0xC5, 0, 0, 0, 0); // mov rax, gs:[rax*8]
Emit(page, ref offset, 0xC3); // ret
return start;
}
// Linux: call pthread_getspecific, preserving the registers that are
// volatile in SysV but non-volatile in Win64 (rsi, rdi).
var pthreadGetSpecific = ResolveLibcExport("pthread_getspecific");
Emit(page, ref offset, 0x56); // push rsi
Emit(page, ref offset, 0x57); // push rdi
Emit(page, ref offset, 0x48, 0x83, 0xEC, 0x08); // sub rsp, 8
Emit(page, ref offset, 0x89, 0xCF); // mov edi, ecx
EmitMovRaxImm64(page, ref offset, pthreadGetSpecific); // mov rax, imm64
Emit(page, ref offset, 0xFF, 0xD0); // call rax
Emit(page, ref offset, 0x48, 0x83, 0xC4, 0x08); // add rsp, 8
Emit(page, ref offset, 0x5F); // pop rdi
Emit(page, ref offset, 0x5E); // pop rsi
Emit(page, ref offset, 0xC3); // ret
return start;
}
private static nint EmitQueryPerformanceCounter(byte* page, ref int offset)
{
// BOOL QueryPerformanceCounter(LARGE_INTEGER* out in rcx): the emitted
// consumers only need a monotonically increasing counter, which rdtsc
// provides without leaving Win64-safe registers.
var start = (nint)(page + offset);
Emit(page, ref offset, 0x0F, 0x31); // rdtsc
Emit(page, ref offset, 0x48, 0xC1, 0xE2, 0x20); // shl rdx, 32
Emit(page, ref offset, 0x48, 0x09, 0xD0); // or rax, rdx
Emit(page, ref offset, 0x48, 0x89, 0x01); // mov [rcx], rax
Emit(page, ref offset, 0xB8, 0x01, 0x00, 0x00, 0x00); // mov eax, 1
Emit(page, ref offset, 0xC3); // ret
return start;
}
private static nint EmitSwitchToThread(byte* page, ref int offset)
{
var schedYield = ResolveLibcExport("sched_yield");
var start = (nint)(page + offset);
Emit(page, ref offset, 0x56); // push rsi
Emit(page, ref offset, 0x57); // push rdi
Emit(page, ref offset, 0x48, 0x83, 0xEC, 0x08); // sub rsp, 8
EmitMovRaxImm64(page, ref offset, schedYield); // mov rax, imm64
Emit(page, ref offset, 0xFF, 0xD0); // call rax
Emit(page, ref offset, 0x48, 0x83, 0xC4, 0x08); // add rsp, 8
Emit(page, ref offset, 0x5F); // pop rdi
Emit(page, ref offset, 0x5E); // pop rsi
Emit(page, ref offset, 0xB8, 0x01, 0x00, 0x00, 0x00); // mov eax, 1
Emit(page, ref offset, 0xC3); // ret
return start;
}
private static nint EmitSleep(byte* page, ref int offset)
{
// void Sleep(DWORD milliseconds in ecx) -> usleep(microseconds in edi).
var usleep = ResolveLibcExport("usleep");
var start = (nint)(page + offset);
Emit(page, ref offset, 0x56); // push rsi
Emit(page, ref offset, 0x57); // push rdi
Emit(page, ref offset, 0x48, 0x83, 0xEC, 0x08); // sub rsp, 8
Emit(page, ref offset, 0x89, 0xCF); // mov edi, ecx
Emit(page, ref offset, 0x81, 0xFF, 0xFF, 0x0F, 0x00, 0x00); // cmp edi, 0xFFF
Emit(page, ref offset, 0x76, 0x05); // jbe +5
Emit(page, ref offset, 0xBF, 0xFF, 0x0F, 0x00, 0x00); // mov edi, 0xFFF (cap at ~4s)
Emit(page, ref offset, 0x69, 0xFF, 0xE8, 0x03, 0x00, 0x00); // imul edi, edi, 1000
EmitMovRaxImm64(page, ref offset, usleep); // mov rax, imm64
Emit(page, ref offset, 0xFF, 0xD0); // call rax
Emit(page, ref offset, 0x48, 0x83, 0xC4, 0x08); // add rsp, 8
Emit(page, ref offset, 0x5F); // pop rdi
Emit(page, ref offset, 0x5E); // pop rsi
Emit(page, ref offset, 0xC3); // ret
return start;
}
private static nint ResolveLibcExport(string name)
{
var libc = NativeLibrary.Load(OperatingSystem.IsMacOS() ? "libSystem.dylib" : "libc.so.6");
return NativeLibrary.GetExport(libc, name);
}
private static void Emit(byte* page, ref int offset, params byte[] bytes)
{
foreach (var value in bytes)
{
page[offset++] = value;
}
}
private static void EmitMovRaxImm64(byte* page, ref int offset, nint value)
{
Emit(page, ref offset, 0x48, 0xB8);
*(long*)(page + offset) = value;
offset += sizeof(long);
}
[DllImport("libc", EntryPoint = "pthread_key_create", SetLastError = true)]
private static extern int pthread_key_create_mac(nuint* key, nint destructor);
[DllImport("libc", EntryPoint = "pthread_key_create", SetLastError = true)]
private static extern int pthread_key_create_linux(uint* key, nint destructor);
[DllImport("libc", EntryPoint = "pthread_key_delete")]
private static extern int pthread_key_delete_mac(nuint key);
[DllImport("libc", EntryPoint = "pthread_key_delete")]
private static extern int pthread_key_delete_linux(uint key);
[DllImport("libc", EntryPoint = "pthread_setspecific")]
private static extern int pthread_setspecific_mac(nuint key, nint value);
[DllImport("libc", EntryPoint = "pthread_setspecific")]
private static extern int pthread_setspecific_linux(uint key, nint value);
[DllImport("libc", EntryPoint = "pthread_getspecific")]
private static extern nint pthread_getspecific_mac(nuint key);
[DllImport("libc", EntryPoint = "pthread_getspecific")]
private static extern nint pthread_getspecific_linux(uint key);
[DllImport("libc")]
private static extern int pthread_threadid_np(nint thread, ulong* threadId);
[DllImport("libc")]
private static extern int gettid();
[DllImport("libc")]
private static extern nint dispatch_semaphore_create(long value);
[DllImport("libc")]
private static extern nint dispatch_semaphore_signal(nint semaphore);
[DllImport("libc")]
private static extern nint dispatch_semaphore_wait(nint semaphore, ulong timeout);
[DllImport("libc")]
private static extern ulong dispatch_time(ulong when, long deltaNanoseconds);
[DllImport("libc")]
private static extern void dispatch_release(nint handle);
[DllImport("libc")]
private static extern int sem_init(nint semaphore, int shared, uint value);
[DllImport("libc")]
private static extern int sem_post(nint semaphore);
[DllImport("libc")]
private static extern int sem_wait(nint semaphore);
[DllImport("libc")]
private static extern int sem_trywait(nint semaphore);
[DllImport("libc")]
private static extern int sem_destroy(nint semaphore);
}
@@ -0,0 +1,115 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using System;
namespace SharpEmu.Core.Cpu.Native;
/// <summary>
/// Recognizes Sony's AMD-only SSE4a EXTRQ+blend idiom and rewrites it into an
/// equivalent SSE4.1 sequence. SharpEmu executes guest x86-64 natively, but
/// Rosetta 2 and Intel hosts do not implement SSE4a, so the original opcode
/// raises #UD -> SIGILL. The compiler emits the idiom against whichever XMM
/// register it happens to allocate (Dead Cells uses xmm1, others xmm2), so the
/// source register is read from the ModRM r/m field rather than hard-coded.
///
/// The match/encode logic is deliberately free of native page-patching so it
/// can be unit-tested against handcrafted byte sequences.
/// </summary>
public static class Sse4aExtrqBlendPatch
{
/// <summary>Length in bytes of both the matched idiom and its replacement.</summary>
public const int SequenceLength = 12;
/// <summary>
/// Matches the 12-byte idiom, extracting the destination register D and the
/// source (scratch) register N:
/// <code>
/// EXTRQ xmmN, 0x28, 0x00 ; 66 0F 78 /0 28 00 mask xmmN to low 40 bits
/// VPBLENDD xmmD, xmmD, xmmN, 2 ; C4 E3 vvvv 02 /r 02 copy dword 1 into xmmD
/// </code>
/// N lives in the ModRM r/m field of both instructions; D (the blend
/// destination and src1) lives in the VPBLENDD ModRM reg field and VEX.vvvv.
/// Both are xmm0-xmm7 (the VEX byte1 0xE3 pins R/X/B, so no xmm8-15 extension).
/// The compiler allocates whichever registers it likes — Dead Cells builds use
/// D=xmm0 and D=xmm3, others differ — so both are read from the encoding.
/// </summary>
public static bool TryMatch(ReadOnlySpan<byte> source, out int destRegister, out int srcRegister)
{
destRegister = -1;
srcRegister = -1;
if (source.Length < SequenceLength)
{
return false;
}
// EXTRQ xmmN, 0x28, 0x00 : 66 0F 78, ModRM (mod=11 reg=000 rm=N), 28, 00.
if (source[0] != 0x66 || source[1] != 0x0F || source[2] != 0x78 ||
(source[3] & 0xF8) != 0xC0 || source[4] != 0x28 || source[5] != 0x00)
{
return false;
}
var n = source[3] & 0x07;
// VPBLENDD xmmD, xmmD, xmmN, 2 : C4 E3 <W=0 vvvv=~D L=0 pp=01> 02 ModRM 02.
// VEX.byte2 fixed bits (W, L, pp) must read 0b*0000*01; vvvv encodes ~D.
if (source[6] != 0xC4 || source[7] != 0xE3 || (source[8] & 0x87) != 0x01 ||
source[9] != 0x02 || source[11] != 0x02)
{
return false;
}
var d = (~(source[8] >> 3)) & 0x0F;
if (d > 7)
{
return false;
}
// ModRM: mod=11, reg=D (dest = src1), rm=N (src2 = the masked register).
if (source[10] != (0xC0 | (d << 3) | n))
{
return false;
}
destRegister = d;
srcRegister = n;
return true;
}
/// <summary>
/// Writes the SSE4.1 equivalent into <paramref name="destination"/>:
/// <code>
/// PEXTRB eax, xmmN, 4 ; 66 0F 3A 14 /r 04 extract byte 4 (zero-extended)
/// PINSRD xmmD, eax, 1 ; 66 0F 3A 22 /r 01 insert into xmmD dword lane 1
/// </code>
/// After EXTRQ masks xmmN to its low 40 bits, dword 1 is just byte 4
/// zero-extended, so the two-instruction extract/insert reproduces the exact
/// observable result the AMD idiom left in xmmD. eax is a caller-dead scratch
/// at every site the compiler emits this idiom.
/// </summary>
public static bool TryEncode(int destRegister, int srcRegister, Span<byte> destination)
{
if ((uint)destRegister > 7 || (uint)srcRegister > 7 || destination.Length < SequenceLength)
{
return false;
}
// PEXTRB eax, xmmN, 4 : ModRM (mod=11 reg=N rm=000 -> eax), imm8 = byte index 4.
destination[0] = 0x66;
destination[1] = 0x0F;
destination[2] = 0x3A;
destination[3] = 0x14;
destination[4] = (byte)(0xC0 | (srcRegister << 3));
destination[5] = 0x04;
// PINSRD xmmD, eax, 1 : ModRM (mod=11 reg=D -> xmmD, rm=000 -> eax), lane 1.
destination[6] = 0x66;
destination[7] = 0x0F;
destination[8] = 0x3A;
destination[9] = 0x22;
destination[10] = (byte)(0xC0 | (destRegister << 3));
destination[11] = 0x01;
return true;
}
}
+4 -4
View File
@@ -194,11 +194,11 @@ public sealed unsafe class StubManager : IDisposable
_stubAddresses.Clear();
}
[DllImport("kernel32.dll", SetLastError = true)]
private static extern void* VirtualAlloc(void* lpAddress, nuint dwSize, AllocationType flAllocationType, MemoryProtection flProtect);
private static void* VirtualAlloc(void* lpAddress, nuint dwSize, AllocationType flAllocationType, MemoryProtection flProtect) =>
HostMemory.Alloc(lpAddress, dwSize, (uint)flAllocationType, (uint)flProtect);
[DllImport("kernel32.dll", SetLastError = true)]
private static extern bool VirtualFree(void* lpAddress, nuint dwSize, FreeType dwFreeType);
private static bool VirtualFree(void* lpAddress, nuint dwSize, FreeType dwFreeType) =>
HostMemory.Free(lpAddress, dwSize, (uint)dwFreeType);
[Flags]
private enum AllocationType : uint
@@ -0,0 +1,33 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
namespace SharpEmu.Core.Cpu.Native.Windows;
/// <summary>
/// Byte offsets into the Win64 CONTEXT record delivered to vectored exception
/// handlers. The handlers read/write guest registers directly at these offsets
/// (no managed CONTEXT struct exists); a future POSIX backend gets a sibling
/// class for its mcontext layout.
/// </summary>
internal static class Win64ContextOffsets
{
public const int Size = 0x4D0;
public const int Mxcsr = 52;
public const int Rax = 120;
public const int Rcx = 128;
public const int Rdx = 136;
public const int Rbx = 144;
public const int Rsp = 152;
public const int Rbp = 160;
public const int Rsi = 168;
public const int Rdi = 176;
public const int R8 = 184;
public const int R9 = 192;
public const int R10 = 200;
public const int R11 = 208;
public const int R12 = 216;
public const int R13 = 224;
public const int R14 = 232;
public const int R15 = 240;
public const int Rip = 248;
}
@@ -0,0 +1,24 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
namespace SharpEmu.Core.Cpu.Native.Windows;
/// <summary>
/// Windows NTSTATUS exception codes and EXCEPTION_RECORD access-type values the
/// fault handlers filter on. Values are the same numbers the handlers previously
/// compared as bare literals; only the spelling changed.
/// </summary>
internal static class WindowsFaultCodes
{
public const uint AccessViolation = 0xC0000005u; // 3221225477
public const uint Breakpoint = 0x80000003u; // 2147483651
public const uint IllegalInstruction = 0xC000001Du; // 3221225501
public const uint FastFail = 0xC0000409u; // 3221226505
public const uint StackOverflow = 0xC00000FDu;
public const uint ClrManagedException = 0xE0434352u;
// EXCEPTION_RECORD.ExceptionInformation[0] for access violations.
public const ulong AccessRead = 0;
public const ulong AccessWrite = 1;
public const ulong AccessExecute = 8;
}
@@ -0,0 +1,191 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using System.Runtime.InteropServices;
using SharpEmu.HLE.Host;
namespace SharpEmu.Core.Cpu.Native.Windows;
/// <summary>
/// Vectored-exception-handler installation and the handler pre-filter thunk.
/// The thunk is inherently Windows-shaped (TEB stack-limit reads via gs:,
/// NTSTATUS pre-filtering, Win64 calling convention) and moved here whole from
/// DirectExecutionBackend; a POSIX backend supplies a sibling built around
/// sigaction/sigaltstack instead.
/// </summary>
internal sealed unsafe partial class WindowsFaultHandling : IHostFaultHandling
{
private readonly IHostMemory _memory;
public WindowsFaultHandling(IHostMemory memory)
{
_memory = memory;
}
public nint CreateHandlerThunk(nint managedCallback, uint hostRspSwitchTlsSlot, nint tlsGetValueAddress)
{
const uint stubSize = 256u;
void* ptr = (void*)_memory.Allocate(0, stubSize, HostPageProtection.ReadWriteExecute);
if (ptr == null)
{
return 0;
}
byte* code = (byte*)ptr;
int offset = 0;
// Native pre-filter: these exception codes are raised while the thread can be in
// cooperative GC mode (a C# throw is RaiseException(0xE0434352) on the throwing
// thread; FailFast/stack-overflow arrive mid-runtime-failure). Entering the managed
// handler then trips the CLR's reverse-P/Invoke check and kills the process with
// "Invalid Program: attempted to call a UnmanagedCallersOnly method from managed
// code" — this is why no managed throw (even one with a catch handler) ever
// survived inside the emulator. Continue the handler search without touching
// managed code; the CLR's own VEH handles its exceptions. MSVC C++ exceptions
// (Vulkan drivers, host CRT) are excluded too: the managed handler only ever
// returned CONTINUE_SEARCH for them.
ReadOnlySpan<uint> nonManagedExceptionCodes =
[WindowsFaultCodes.ClrManagedException, 0xE06D7363u, WindowsFaultCodes.FastFail, WindowsFaultCodes.StackOverflow];
EmitByte(code, ref offset, 0x48); EmitByte(code, ref offset, 0x8B); EmitByte(code, ref offset, 0x01); // mov rax, [rcx] (ExceptionRecord*)
EmitByte(code, ref offset, 0x8B); EmitByte(code, ref offset, 0x00); // mov eax, [rax] (ExceptionCode)
var passJumpOffsets = stackalloc int[nonManagedExceptionCodes.Length];
for (int i = 0; i < nonManagedExceptionCodes.Length; i++)
{
EmitByte(code, ref offset, 0x3D); // cmp eax, imm32
EmitUInt32(code, ref offset, nonManagedExceptionCodes[i]);
EmitByte(code, ref offset, 0x74); // je pass
passJumpOffsets[i] = offset;
EmitByte(code, ref offset, 0x00);
}
EmitByte(code, ref offset, 0xEB); EmitByte(code, ref offset, 0x03); // jmp over pass block
int passOffset = offset;
EmitByte(code, ref offset, 0x31); EmitByte(code, ref offset, 0xC0); // pass: xor eax, eax (EXCEPTION_CONTINUE_SEARCH)
EmitByte(code, ref offset, 0xC3); // ret
for (int i = 0; i < nonManagedExceptionCodes.Length; i++)
{
code[passJumpOffsets[i]] = checked((byte)(passOffset - (passJumpOffsets[i] + 1)));
}
EmitByte(code, ref offset, 0x41); EmitByte(code, ref offset, 0x54); // push r12
EmitByte(code, ref offset, 0x41); EmitByte(code, ref offset, 0x55); // push r13
EmitByte(code, ref offset, 0x49); EmitByte(code, ref offset, 0x89); EmitByte(code, ref offset, 0xE4); // mov r12, rsp
EmitByte(code, ref offset, 0x49); EmitByte(code, ref offset, 0x89); EmitByte(code, ref offset, 0xCD); // mov r13, rcx
EmitByte(code, ref offset, 0x65); EmitByte(code, ref offset, 0x48); // mov rax, gs:[8]
EmitByte(code, ref offset, 0x8B); EmitByte(code, ref offset, 0x04); EmitByte(code, ref offset, 0x25);
EmitUInt32(code, ref offset, 8u);
EmitByte(code, ref offset, 0x49); EmitByte(code, ref offset, 0x39); EmitByte(code, ref offset, 0xC4); // cmp r12, rax
EmitByte(code, ref offset, 0x0F); EmitByte(code, ref offset, 0x83); // jae guestStack
int aboveStackJump = offset;
EmitUInt32(code, ref offset, 0u);
EmitByte(code, ref offset, 0x65); EmitByte(code, ref offset, 0x48); // mov rax, gs:[0x10]
EmitByte(code, ref offset, 0x8B); EmitByte(code, ref offset, 0x04); EmitByte(code, ref offset, 0x25);
EmitUInt32(code, ref offset, 0x10u);
EmitByte(code, ref offset, 0x49); EmitByte(code, ref offset, 0x39); EmitByte(code, ref offset, 0xC4); // cmp r12, rax
EmitByte(code, ref offset, 0x0F); EmitByte(code, ref offset, 0x82); // jb guestStack
int belowStackJump = offset;
EmitUInt32(code, ref offset, 0u);
EmitByte(code, ref offset, 0x48); EmitByte(code, ref offset, 0x83); EmitByte(code, ref offset, 0xEC); EmitByte(code, ref offset, 0x28);
EmitByte(code, ref offset, 0x4C); EmitByte(code, ref offset, 0x89); EmitByte(code, ref offset, 0xE9); // mov rcx, r13
EmitByte(code, ref offset, 0x48); EmitByte(code, ref offset, 0xB8);
*(nint*)(code + offset) = managedCallback;
offset += sizeof(nint);
EmitByte(code, ref offset, 0xFF); EmitByte(code, ref offset, 0xD0);
EmitByte(code, ref offset, 0x48); EmitByte(code, ref offset, 0x83); EmitByte(code, ref offset, 0xC4); EmitByte(code, ref offset, 0x28);
EmitByte(code, ref offset, 0xE9);
int hostRestoreJump = offset;
EmitUInt32(code, ref offset, 0u);
int guestStackOffset = offset;
EmitByte(code, ref offset, 0x48); EmitByte(code, ref offset, 0x83); EmitByte(code, ref offset, 0xEC); EmitByte(code, ref offset, 0x28);
EmitByte(code, ref offset, 0xB9);
EmitUInt32(code, ref offset, hostRspSwitchTlsSlot);
EmitByte(code, ref offset, 0x48); EmitByte(code, ref offset, 0xB8);
*(nint*)(code + offset) = tlsGetValueAddress;
offset += sizeof(nint);
EmitByte(code, ref offset, 0xFF); EmitByte(code, ref offset, 0xD0);
EmitByte(code, ref offset, 0x48); EmitByte(code, ref offset, 0x83); EmitByte(code, ref offset, 0xC4); EmitByte(code, ref offset, 0x28);
EmitByte(code, ref offset, 0x48); EmitByte(code, ref offset, 0x85); EmitByte(code, ref offset, 0xC0); // test rax, rax
EmitByte(code, ref offset, 0x0F); EmitByte(code, ref offset, 0x84);
int missingTlsJump = offset;
EmitUInt32(code, ref offset, 0u);
EmitByte(code, ref offset, 0x4C); EmitByte(code, ref offset, 0x8B); EmitByte(code, ref offset, 0x18); // mov r11, [rax]
EmitByte(code, ref offset, 0x4D); EmitByte(code, ref offset, 0x85); EmitByte(code, ref offset, 0xDB); // test r11, r11
EmitByte(code, ref offset, 0x0F); EmitByte(code, ref offset, 0x84);
int missingHostStackJump = offset;
EmitUInt32(code, ref offset, 0u);
EmitByte(code, ref offset, 0x4C); EmitByte(code, ref offset, 0x89); EmitByte(code, ref offset, 0xDC); // mov rsp, r11
EmitByte(code, ref offset, 0x48); EmitByte(code, ref offset, 0x83); EmitByte(code, ref offset, 0xEC); EmitByte(code, ref offset, 0x28);
EmitByte(code, ref offset, 0x4C); EmitByte(code, ref offset, 0x89); EmitByte(code, ref offset, 0xE9); // mov rcx, r13
EmitByte(code, ref offset, 0x48); EmitByte(code, ref offset, 0xB8);
*(nint*)(code + offset) = managedCallback;
offset += sizeof(nint);
EmitByte(code, ref offset, 0xFF); EmitByte(code, ref offset, 0xD0);
EmitByte(code, ref offset, 0x48); EmitByte(code, ref offset, 0x83); EmitByte(code, ref offset, 0xC4); EmitByte(code, ref offset, 0x28);
EmitByte(code, ref offset, 0xE9);
int guestRestoreJump = offset;
EmitUInt32(code, ref offset, 0u);
int passThroughOffset = offset;
EmitByte(code, ref offset, 0x31); EmitByte(code, ref offset, 0xC0); // xor eax, eax
int restoreOffset = offset;
EmitByte(code, ref offset, 0x4C); EmitByte(code, ref offset, 0x89); EmitByte(code, ref offset, 0xE4); // mov rsp, r12
EmitByte(code, ref offset, 0x41); EmitByte(code, ref offset, 0x5D);
EmitByte(code, ref offset, 0x41); EmitByte(code, ref offset, 0x5C);
EmitByte(code, ref offset, 0xC3);
*(int*)(code + aboveStackJump) = guestStackOffset - (aboveStackJump + sizeof(int));
*(int*)(code + belowStackJump) = guestStackOffset - (belowStackJump + sizeof(int));
*(int*)(code + hostRestoreJump) = restoreOffset - (hostRestoreJump + sizeof(int));
*(int*)(code + missingTlsJump) = passThroughOffset - (missingTlsJump + sizeof(int));
*(int*)(code + missingHostStackJump) = passThroughOffset - (missingHostStackJump + sizeof(int));
*(int*)(code + guestRestoreJump) = restoreOffset - (guestRestoreJump + sizeof(int));
if (!_memory.Protect((ulong)ptr, stubSize, HostPageProtection.ReadExecute, out _))
{
Console.Error.WriteLine($"[LOADER][ERROR] VirtualProtect failed for exception handler trampoline at 0x{(nint)ptr:X16}");
_ = _memory.Free((ulong)ptr);
return 0;
}
_memory.FlushInstructionCache((ulong)ptr, (ulong)offset);
return (nint)ptr;
}
public void FreeThunk(nint thunk)
{
_ = _memory.Free((ulong)thunk);
}
public nint AddFirstChanceHandler(nint thunk)
{
return (nint)AddVectoredExceptionHandler(1u, thunk);
}
public void RemoveHandler(nint handle)
{
_ = RemoveVectoredExceptionHandler((void*)handle);
}
public void SetUnhandledFilter(nint thunk)
{
_ = SetUnhandledExceptionFilter(thunk);
}
private static void EmitByte(byte* code, ref int offset, byte value)
{
code[offset++] = value;
}
private static void EmitUInt32(byte* code, ref int offset, uint value)
{
*(uint*)(code + offset) = value;
offset += sizeof(uint);
}
[LibraryImport("kernel32.dll")]
private static partial void* AddVectoredExceptionHandler(uint first, IntPtr handler);
[LibraryImport("kernel32.dll")]
private static partial uint RemoveVectoredExceptionHandler(void* handle);
[LibraryImport("kernel32.dll")]
private static partial IntPtr SetUnhandledExceptionFilter(IntPtr lpTopLevelExceptionFilter);
}
+9 -1
View File
@@ -5,7 +5,7 @@ using SharpEmu.HLE;
namespace SharpEmu.Core.Cpu;
public sealed class TrackedCpuMemory : ICpuMemory, ITrackedCpuMemory, IGuestMemoryAllocator
public sealed class TrackedCpuMemory : ICpuMemory, ITrackedCpuMemory, IGuestMemoryAllocator, ICpuMemoryWrapper
{
private readonly ICpuMemory _inner;
@@ -40,6 +40,9 @@ public sealed class TrackedCpuMemory : ICpuMemory, ITrackedCpuMemory, IGuestMemo
return result;
}
public bool TryCopy(ulong destinationAddress, ulong sourceAddress, ulong length) =>
_inner.TryCopy(destinationAddress, sourceAddress, length);
public bool TryAllocateGuestMemory(ulong size, ulong alignment, out ulong address)
{
if (_inner is IGuestMemoryAllocator allocator)
@@ -50,4 +53,9 @@ public sealed class TrackedCpuMemory : ICpuMemory, ITrackedCpuMemory, IGuestMemo
address = 0;
return false;
}
public bool TryFreeGuestMemory(ulong address)
{
return _inner is IGuestMemoryAllocator allocator && allocator.TryFreeGuestMemory(address);
}
}
@@ -48,6 +48,8 @@ public enum ProgramHeaderType : uint
Tls = 7,
GnuEhFrame = 0x6474E550,
SceRela = 0x60000000,
SceProcParam = 0x61000001,
+14 -1
View File
@@ -24,7 +24,10 @@ public sealed class SelfImage
ulong procParamAddress = 0,
string? title = null,
string? titleId = null,
string? version = null)
string? version = null,
uint tlsModuleId = 0,
ulong tlsMemorySize = 0,
ulong tlsStaticOffset = 0)
{
ArgumentNullException.ThrowIfNull(programHeaders);
ArgumentNullException.ThrowIfNull(mappedRegions);
@@ -44,6 +47,9 @@ public sealed class SelfImage
Title = title;
TitleId = titleId;
Version = version;
TlsModuleId = tlsModuleId;
TlsMemorySize = tlsMemorySize;
TlsStaticOffset = tlsStaticOffset;
}
public bool IsSelf { get; }
@@ -75,4 +81,11 @@ public sealed class SelfImage
public string? TitleId { get; }
public string? Version { get; }
public uint TlsModuleId { get; }
public ulong TlsMemorySize { get; }
/// <summary>Variant II distance from the thread pointer to this module's static TLS base.</summary>
public ulong TlsStaticOffset { get; }
}
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+124 -30
View File
@@ -2,6 +2,7 @@
// SPDX-License-Identifier: GPL-2.0-or-later
using SharpEmu.Core.Loader;
using SharpEmu.HLE;
namespace SharpEmu.Core.Memory;
@@ -41,15 +42,14 @@ public sealed class VirtualMemory : IVirtualMemory
lock (_gate)
{
foreach (var existing in _regions)
var insertionIndex = FindInsertionIndex(virtualAddress);
if ((insertionIndex > 0 && virtualAddress < _regions[insertionIndex - 1].EndAddress) ||
(insertionIndex < _regions.Count && endAddress > _regions[insertionIndex].Region.VirtualAddress))
{
if (virtualAddress < existing.EndAddress && endAddress > existing.Region.VirtualAddress)
{
throw new InvalidOperationException("Attempted to map an overlapping virtual memory region.");
}
throw new InvalidOperationException("Attempted to map an overlapping virtual memory region.");
}
_regions.Add(new MappedRegion(
_regions.Insert(insertionIndex, new MappedRegion(
new VirtualMemoryRegion(virtualAddress, memorySize, fileOffset, (ulong)fileData.Length, protection),
endAddress,
backingMemory));
@@ -74,12 +74,12 @@ public sealed class VirtualMemory : IVirtualMemory
{
lock (_gate)
{
if (!TryResolveRegion(virtualAddress, destination.Length, out var region, out var offset))
if (!TryValidateRange(virtualAddress, destination.Length, ProgramHeaderFlags.Read, out var regionIndex))
{
return false;
}
region.BackingMemory.AsSpan(offset, destination.Length).CopyTo(destination);
CopyFromRegions(virtualAddress, destination, regionIndex);
return true;
}
}
@@ -88,39 +88,133 @@ public sealed class VirtualMemory : IVirtualMemory
{
lock (_gate)
{
if (!TryResolveRegion(virtualAddress, source.Length, out var region, out var offset))
if (!TryValidateRange(virtualAddress, source.Length, ProgramHeaderFlags.Write, out var regionIndex))
{
return false;
}
source.CopyTo(region.BackingMemory.AsSpan(offset, source.Length));
return true;
CopyToRegions(virtualAddress, source, regionIndex);
}
if (GuestWriteWatch.Armed)
{
GuestWriteWatch.Check(virtualAddress, source);
}
return true;
}
private bool TryValidateRange(
ulong virtualAddress,
int length,
ProgramHeaderFlags requiredProtection,
out int regionIndex)
{
regionIndex = FindContainingRegionIndex(virtualAddress);
if (regionIndex < 0)
{
return false;
}
var currentAddress = virtualAddress;
var remaining = length;
var currentIndex = regionIndex;
while (true)
{
if (currentIndex >= _regions.Count)
{
return false;
}
var region = _regions[currentIndex];
if (currentAddress < region.Region.VirtualAddress ||
currentAddress >= region.EndAddress ||
(region.Region.Protection & requiredProtection) == 0)
{
return false;
}
if (remaining == 0)
{
return true;
}
var available = region.EndAddress - currentAddress;
var chunkLength = (int)Math.Min((ulong)remaining, available);
remaining -= chunkLength;
if (remaining == 0)
{
return true;
}
currentAddress += (ulong)chunkLength;
currentIndex++;
}
}
private bool TryResolveRegion(ulong virtualAddress, int length, out MappedRegion region, out int offset)
private int FindContainingRegionIndex(ulong virtualAddress)
{
foreach (var candidate in _regions)
var insertionIndex = FindInsertionIndex(virtualAddress);
if (insertionIndex < _regions.Count &&
_regions[insertionIndex].Region.VirtualAddress == virtualAddress)
{
if (virtualAddress < candidate.Region.VirtualAddress || virtualAddress >= candidate.EndAddress)
{
continue;
}
var candidateOffset = checked((int)(virtualAddress - candidate.Region.VirtualAddress));
if (candidateOffset + length > candidate.BackingMemory.Length)
{
break;
}
region = candidate;
offset = candidateOffset;
return true;
return insertionIndex;
}
region = default;
offset = 0;
return false;
var candidateIndex = insertionIndex - 1;
return candidateIndex >= 0 && virtualAddress < _regions[candidateIndex].EndAddress
? candidateIndex
: -1;
}
private void CopyFromRegions(ulong virtualAddress, Span<byte> destination, int regionIndex)
{
var copied = 0;
var currentAddress = virtualAddress;
while (copied < destination.Length)
{
var region = _regions[regionIndex++];
var regionOffset = checked((int)(currentAddress - region.Region.VirtualAddress));
var chunkLength = Math.Min(destination.Length - copied, region.BackingMemory.Length - regionOffset);
region.BackingMemory.AsSpan(regionOffset, chunkLength).CopyTo(destination[copied..]);
copied += chunkLength;
currentAddress += (ulong)chunkLength;
}
}
private void CopyToRegions(ulong virtualAddress, ReadOnlySpan<byte> source, int regionIndex)
{
var copied = 0;
var currentAddress = virtualAddress;
while (copied < source.Length)
{
var region = _regions[regionIndex++];
var regionOffset = checked((int)(currentAddress - region.Region.VirtualAddress));
var chunkLength = Math.Min(source.Length - copied, region.BackingMemory.Length - regionOffset);
source.Slice(copied, chunkLength).CopyTo(region.BackingMemory.AsSpan(regionOffset, chunkLength));
copied += chunkLength;
currentAddress += (ulong)chunkLength;
}
}
private int FindInsertionIndex(ulong virtualAddress)
{
var lower = 0;
var upper = _regions.Count;
while (lower < upper)
{
var middle = lower + ((upper - lower) / 2);
if (_regions[middle].Region.VirtualAddress < virtualAddress)
{
lower = middle + 1;
}
else
{
upper = middle;
}
}
return lower;
}
private readonly record struct MappedRegion(VirtualMemoryRegion Region, ulong EndAddress, byte[] BackingMemory);
+208 -121
View File
@@ -11,20 +11,18 @@ using SharpEmu.Libs.VideoOut;
using SharpEmu.Libs.Kernel;
using SharpEmu.Libs.AppContent;
using SharpEmu.Libs.SaveData;
using SharpEmu.Logging;
using SharpEmu.Libs.Fiber;
using SharpEmu.Libs.SystemService;
using System.Buffers.Binary;
using System.Collections.Generic;
using System.Linq;
using System.Text;
using System.IO;
namespace SharpEmu.Core.Runtime;
public sealed class SharpEmuRuntime : ISharpEmuRuntime
{
private static readonly SharpEmuLogger Log = SharpEmuLog.For("Runtime");
private readonly record struct LoadedModuleImage(string Path, SelfImage Image);
private readonly record struct LoadedModuleImage(string Path, SelfImage Image, int Handle, bool StartAtBoot);
private static readonly HashSet<string> PreloadSkipModules = new(StringComparer.OrdinalIgnoreCase)
{
@@ -70,6 +68,7 @@ public sealed class SharpEmuRuntime : ISharpEmuRuntime
CpuEngine = cpuExecutionOptions.CpuEngine,
StrictDynlibResolution = cpuExecutionOptions.StrictDynlibResolution,
ImportTraceLimit = Math.Max(0, cpuExecutionOptions.ImportTraceLimit),
DebugHook = cpuExecutionOptions.DebugHook,
};
_fileSystem = fileSystem ?? new PhysicalFileSystem();
}
@@ -81,9 +80,12 @@ public sealed class SharpEmuRuntime : ISharpEmuRuntime
CpuEngine = options.CpuEngine,
StrictDynlibResolution = options.StrictDynlibResolution,
ImportTraceLimit = Math.Max(0, options.ImportTraceLimit),
DebugHook = options.DebugHook,
};
var moduleManager = new ModuleManager();
moduleManager.RegisterFromAssembly(typeof(KernelExports).Assembly, Generation.Gen4 | Generation.Gen5, Aerolib.Instance);
// The compile-time generated registry (SharpEmu.SourceGenerators) is the sole
// registration source; content tests in SharpEmu.Libs.Tests pin its invariants.
moduleManager.RegisterExports(SharpEmu.Generated.SysAbiExportRegistry.CreateExports(Generation.Gen4 | Generation.Gen5));
moduleManager.Freeze();
var virtualMemory = new PhysicalVirtualMemory();
@@ -131,20 +133,21 @@ public sealed class SharpEmuRuntime : ISharpEmuRuntime
{
var normalizedEbootPath = Path.GetFullPath(ebootPath);
using var app0Binding = BindApp0Root(normalizedEbootPath);
Log.Info($"Loading: {ebootPath}");
Console.Error.WriteLine($"[RUNTIME] Loading: {ebootPath}");
LastExecutionDiagnostics = null;
LastExecutionTrace = null;
LastSessionSummary = null;
LastBasicBlockTrace = null;
LastMilestoneLog = null;
FiberExports.ResetRuntimeState();
KernelModuleRegistry.Reset();
var image = LoadImage(normalizedEbootPath);
VideoOutExports.ConfigureApplicationInfo(image.Title, image.TitleId, image.Version, BuildInfo.CommitSha);
VideoOutExports.ConfigureApplicationInfo(image.Title, image.TitleId, image.Version);
SaveDataExports.ConfigureApplicationInfo(image.TitleId);
LogAppBundleInfo(normalizedEbootPath, image);
RegisterLoadedModule(normalizedEbootPath, image, isMain: true, isSystemModule: false);
SystemServiceExports.ConfigureApplicationInfo(image.TitleId);
_ = RegisterLoadedModule(normalizedEbootPath, image, isMain: true, isSystemModule: false);
KernelRuntimeCompatExports.ConfigureProcessProcParamAddress(image.ProcParamAddress);
Log.Info($"Entry: 0x{image.EntryPoint:X16}");
Console.Error.WriteLine($"[RUNTIME] Entry: 0x{image.EntryPoint:X16}");
var generation = image.ElfHeader.AbiVersion == 2 ? Generation.Gen5 : Generation.Gen4;
var activeImportStubs = new Dictionary<ulong, string>(image.ImportStubs);
var activeRuntimeSymbols = new Dictionary<string, ulong>(image.RuntimeSymbols, StringComparer.Ordinal);
@@ -167,7 +170,7 @@ public sealed class SharpEmuRuntime : ISharpEmuRuntime
processImageName);
if (initializerResult is { } failedInitializerResult)
{
Log.Error($"Initializer dispatch failed: {failedInitializerResult}");
Console.Error.WriteLine($"[RUNTIME] Initializer dispatch failed: {failedInitializerResult}");
LastExecutionTrace = _cpuDispatcher.LastImportResolutionTrace;
LastMilestoneLog = _cpuDispatcher.LastMilestoneLog;
LastSessionSummary = BuildSessionSummary(_cpuDispatcher.LastSessionSummary);
@@ -175,8 +178,8 @@ public sealed class SharpEmuRuntime : ISharpEmuRuntime
return failedInitializerResult;
}
Log.Info($"Dispatching, gen: {generation}");
Log.Debug($"About to call DispatchEntry with entryPoint=0x{image.EntryPoint:X16}");
Console.Error.WriteLine($"[RUNTIME] Dispatching, gen: {generation}");
Console.Error.WriteLine($"[RUNTIME] About to call DispatchEntry with entryPoint=0x{image.EntryPoint:X16}");
var result = _cpuDispatcher.DispatchEntry(
image.EntryPoint,
@@ -186,8 +189,18 @@ public sealed class SharpEmuRuntime : ISharpEmuRuntime
processImageName,
_cpuExecutionOptions);
Log.Info($"DispatchEntry returned: {result}");
Log.Info($"Dispatch result: {result}");
Console.Error.WriteLine($"[RUNTIME] DispatchEntry returned: {result}");
Console.Error.WriteLine($"[RUNTIME] Dispatch result: {result}");
// Stop is a host operation, not an emulation failure. The detailed
// trace and session-summary builders can traverse a partially torn
// down native backend, delaying the GUI exit callback indefinitely.
if (HostSessionControl.IsShutdownRequested)
{
Console.Error.WriteLine("[RUNTIME] Skipping post-exit diagnostics for host shutdown.");
return result;
}
LastExecutionTrace = _cpuDispatcher.LastImportResolutionTrace;
LastMilestoneLog = _cpuDispatcher.LastMilestoneLog;
LastSessionSummary = BuildSessionSummary(_cpuDispatcher.LastSessionSummary);
@@ -358,66 +371,6 @@ public sealed class SharpEmuRuntime : ISharpEmuRuntime
return result;
}
private static void LogAppBundleInfo(string ebootPath, SelfImage image)
{
var executableName = Path.GetFileName(ebootPath);
if (string.IsNullOrWhiteSpace(executableName))
{
executableName = "eboot.bin";
}
var displayName = string.IsNullOrWhiteSpace(image.Title) ? "(unknown)" : image.Title!.Trim();
var titleId = string.IsNullOrWhiteSpace(image.TitleId) ? "(unknown)" : image.TitleId!.Trim();
var version = string.IsNullOrWhiteSpace(image.Version) ? "(unknown)" : image.Version!.Trim();
var contentId = ResolveContentId(Path.GetDirectoryName(ebootPath)) ?? "(unknown)";
var builder = new StringBuilder();
builder.AppendLine("App bundle info:");
builder.AppendLine($"- Display name: {displayName}");
builder.AppendLine($"- Version: {version}");
builder.AppendLine($"- Title ID: {titleId}");
builder.AppendLine($"- Content ID: {contentId}");
builder.AppendLine($"- Executable: {executableName}");
builder.Append("- Platform: PlayStation 5");
Log.Info(builder.ToString());
}
private static string? ResolveContentId(string? bundleRoot)
{
if (string.IsNullOrEmpty(bundleRoot))
{
return null;
}
foreach (var candidate in new[]
{
Path.Combine(bundleRoot, "sce_sys", "param.json"),
Path.Combine(bundleRoot, "param.json"),
})
{
if (!File.Exists(candidate))
{
continue;
}
try
{
using var doc = System.Text.Json.JsonDocument.Parse(File.ReadAllBytes(candidate));
if (doc.RootElement.TryGetProperty("contentId", out var contentId))
{
return contentId.GetString();
}
}
catch (Exception ex) when (ex is IOException or System.Text.Json.JsonException)
{
}
break;
}
return null;
}
private static App0BindingScope? BindApp0Root(string normalizedEbootPath)
{
const string app0VariableName = "SHARPEMU_APP0_DIR";
@@ -432,6 +385,20 @@ public sealed class SharpEmuRuntime : ISharpEmuRuntime
return null;
}
// Dump tools commonly place decrypted executables in an app0/decrypted
// sidecar while leaving Unity data in the parent app0 directory. Keep
// loading code/modules beside the decrypted eboot, but mount /app0 at
// the content-bearing parent so boot.config, metadata and assets resolve.
if (string.Equals(Path.GetFileName(app0Root), "decrypted", StringComparison.OrdinalIgnoreCase))
{
var parentRoot = Path.GetDirectoryName(app0Root);
if (!string.IsNullOrWhiteSpace(parentRoot) &&
Directory.Exists(Path.Combine(parentRoot, "Media")))
{
app0Root = parentRoot;
}
}
Environment.SetEnvironmentVariable(app0VariableName, app0Root);
return new App0BindingScope(app0VariableName);
}
@@ -477,6 +444,56 @@ public sealed class SharpEmuRuntime : ISharpEmuRuntime
return null;
}
private bool TryGetEhFrameInfo(
SelfImage image,
ulong imageSize,
out ulong ehFrameHeaderAddress,
out ulong ehFrameAddress,
out ulong ehFrameSize)
{
ehFrameHeaderAddress = 0;
ehFrameAddress = 0;
ehFrameSize = 0;
var imageBase = image.EntryPoint >= image.ElfHeader.EntryPoint
? image.EntryPoint - image.ElfHeader.EntryPoint
: 0UL;
for (var i = 0; i < image.ProgramHeaders.Count; i++)
{
var header = image.ProgramHeaders[i];
if (header.HeaderType != ProgramHeaderType.GnuEhFrame || header.MemorySize < 8)
{
continue;
}
var headerAddress = imageBase + header.VirtualAddress;
Span<byte> ehHeader = stackalloc byte[8];
if (!_virtualMemory.TryRead(headerAddress, ehHeader) ||
ehHeader[0] != 1 ||
ehHeader[1] != 0x1B)
{
continue;
}
var relativeOffset = BinaryPrimitives.ReadInt32LittleEndian(ehHeader[4..]);
ehFrameHeaderAddress = headerAddress;
ehFrameAddress = unchecked((ulong)((long)headerAddress + 4 + relativeOffset));
if (ehFrameAddress < 0x10000)
{
continue;
}
var relativeEhFrameAddress = ehFrameAddress - imageBase;
ehFrameSize = header.VirtualAddress > relativeEhFrameAddress
? header.VirtualAddress - relativeEhFrameAddress
: imageSize > header.VirtualAddress
? imageSize - header.VirtualAddress
: 0;
return true;
}
return false;
}
private OrbisGen2Result? RunPreloadedModuleInitializers(
IReadOnlyList<LoadedModuleImage> loadedModuleImages,
Generation generation,
@@ -486,20 +503,30 @@ public sealed class SharpEmuRuntime : ISharpEmuRuntime
for (var i = 0; i < loadedModuleImages.Count; i++)
{
var loadedModule = loadedModuleImages[i];
if (!loadedModule.StartAtBoot)
{
continue;
}
var initEntryPoint = loadedModule.Image.InitFunctionEntryPoint;
if (initEntryPoint < 0x10000)
{
continue;
}
if (!KernelModuleRegistry.TryBeginModuleStart(loadedModule.Handle, out _))
{
continue;
}
var moduleName = Path.GetFileName(loadedModule.Path);
if (string.IsNullOrWhiteSpace(moduleName))
{
moduleName = $"module#{i}";
}
Log.Info(
$"Starting module {moduleName}: dt_init=0x{initEntryPoint:X16}");
Console.Error.WriteLine(
$"[RUNTIME] Starting module {moduleName}: dt_init=0x{initEntryPoint:X16}");
var result = _cpuDispatcher.DispatchModuleInitializer(
initEntryPoint,
@@ -508,10 +535,13 @@ public sealed class SharpEmuRuntime : ISharpEmuRuntime
activeRuntimeSymbols,
moduleName,
_cpuExecutionOptions);
KernelModuleRegistry.CompleteModuleStart(
loadedModule.Handle,
result == OrbisGen2Result.ORBIS_GEN2_OK);
if (result != OrbisGen2Result.ORBIS_GEN2_OK)
{
Log.Error(
$"Module start failed: {moduleName} -> {result}");
Console.Error.WriteLine(
$"[RUNTIME] Module start failed: {moduleName} -> {result}");
return result;
}
}
@@ -532,8 +562,8 @@ public sealed class SharpEmuRuntime : ISharpEmuRuntime
return null;
}
Log.Info(
$"Running initializers for {label}: preinit={image.PreInitializerFunctions.Count}, init={image.InitializerFunctions.Count}");
Console.Error.WriteLine(
$"[RUNTIME] Running initializers for {label}: preinit={image.PreInitializerFunctions.Count}, init={image.InitializerFunctions.Count}");
var result = RunInitializerList(
$"{label}:preinit",
@@ -572,8 +602,8 @@ public sealed class SharpEmuRuntime : ISharpEmuRuntime
continue;
}
Log.Debug(
$" Initializer {label}[{i}] -> 0x{initializerAddress:X16}");
Console.Error.WriteLine(
$"[RUNTIME] Initializer {label}[{i}] -> 0x{initializerAddress:X16}");
var result = _cpuDispatcher.DispatchEntry(
initializerAddress,
@@ -605,12 +635,17 @@ public sealed class SharpEmuRuntime : ISharpEmuRuntime
var moduleDirectories = new[]
{
Path.Combine(ebootDirectory, "sce_module"),
Path.Combine(ebootDirectory, "sce_modules"),
Path.Combine(ebootDirectory, "Media", "Modules"),
(Path: Path.Combine(ebootDirectory, "sce_module"), StartAtBoot: true),
(Path: Path.Combine(ebootDirectory, "sce_modules"), StartAtBoot: true),
(Path: Path.Combine(ebootDirectory, "Media", "Modules"), StartAtBoot: true),
// Unity native plugins are loaded later through sceKernelLoadStartModule. Map
// them up front so the HLE loader can return a real module handle and dlsym
// can resolve their exports, but defer DT_INIT until the guest requests them.
(Path: Path.Combine(ebootDirectory, "Media", "Plugins"), StartAtBoot: false),
}
.Distinct(StringComparer.OrdinalIgnoreCase)
.Where(Directory.Exists)
.GroupBy(entry => entry.Path, StringComparer.OrdinalIgnoreCase)
.Select(group => group.First())
.Where(entry => Directory.Exists(entry.Path))
.ToArray();
if (moduleDirectories.Length == 0)
@@ -620,28 +655,30 @@ public sealed class SharpEmuRuntime : ISharpEmuRuntime
var allModulePaths = moduleDirectories
.SelectMany(directory => Directory
.EnumerateFiles(directory)
.OrderBy(path => path, StringComparer.OrdinalIgnoreCase))
.Where(path =>
.EnumerateFiles(directory.Path)
.OrderBy(path => path, StringComparer.OrdinalIgnoreCase)
.Select(path => (Path: path, directory.StartAtBoot)))
.Where(entry =>
{
var extension = Path.GetExtension(path);
var extension = Path.GetExtension(entry.Path);
return string.Equals(extension, ".prx", StringComparison.OrdinalIgnoreCase) ||
string.Equals(extension, ".sprx", StringComparison.OrdinalIgnoreCase);
})
.Distinct(StringComparer.OrdinalIgnoreCase)
.GroupBy(entry => entry.Path, StringComparer.OrdinalIgnoreCase)
.Select(group => group.First())
.ToArray();
var modulePaths = allModulePaths
.Where(ShouldPreloadModule)
.Where(entry => ShouldPreloadModule(entry.Path))
.ToArray();
var skippedModules = allModulePaths
.Where(path => !ShouldPreloadModule(path))
.Select(Path.GetFileName)
.Where(entry => !ShouldPreloadModule(entry.Path))
.Select(entry => Path.GetFileName(entry.Path))
.Where(name => !string.IsNullOrWhiteSpace(name))
.ToArray();
if (skippedModules.Length > 0)
{
Log.Info($"Skipping {skippedModules.Length} core module(s): {string.Join(", ", skippedModules)}");
Console.Error.WriteLine($"[RUNTIME] Skipping {skippedModules.Length} core module(s): {string.Join(", ", skippedModules)}");
}
if (modulePaths.Length == 0)
@@ -649,14 +686,15 @@ public sealed class SharpEmuRuntime : ISharpEmuRuntime
return loadedImages;
}
Log.Debug($"Module search directories: {string.Join(", ", moduleDirectories)}");
Log.Info($"Loading {modulePaths.Length} module(s)...");
Console.Error.WriteLine($"[RUNTIME] Module search directories: {string.Join(", ", moduleDirectories.Select(entry => entry.Path))}");
Console.Error.WriteLine($"[RUNTIME] Loading {modulePaths.Length} module(s)...");
var loadedModules = 0;
var failedModules = 0;
var mergedImportCount = 0;
var mergedSymbolCount = 0;
foreach (var modulePath in modulePaths)
foreach (var moduleEntry in modulePaths)
{
var modulePath = moduleEntry.Path;
try
{
var fileInfo = new FileInfo(modulePath);
@@ -681,25 +719,52 @@ public sealed class SharpEmuRuntime : ISharpEmuRuntime
mergedImportCount += MergeImportStubs(importStubs, moduleImage.ImportStubs, modulePath);
mergedSymbolCount += MergeRuntimeSymbols(runtimeSymbols, moduleImage.RuntimeSymbols);
RegisterLoadedModule(modulePath, moduleImage, isMain: false, isSystemModule: false);
loadedImages.Add(new LoadedModuleImage(modulePath, moduleImage));
InstallNativePluginCompatibilityHooks(importStubs, moduleImage, modulePath);
var moduleHandle = RegisterLoadedModule(modulePath, moduleImage, isMain: false, isSystemModule: false);
var moduleName = Path.GetFileName(modulePath);
var startAtBoot = moduleEntry.StartAtBoot ||
string.Equals(moduleName, "libfmod.prx", StringComparison.OrdinalIgnoreCase) ||
string.Equals(moduleName, "libfmodstudio.prx", StringComparison.OrdinalIgnoreCase);
loadedImages.Add(new LoadedModuleImage(modulePath, moduleImage, moduleHandle, startAtBoot));
loadedModules++;
Log.Info(
$"Loaded module {Path.GetFileName(modulePath)}: entry=0x{moduleImage.EntryPoint:X16}, imports={moduleImage.ImportStubs.Count}, symbols={moduleImage.RuntimeSymbols.Count}");
Console.Error.WriteLine(
$"[RUNTIME] Loaded module {Path.GetFileName(modulePath)}: entry=0x{moduleImage.EntryPoint:X16}, imports={moduleImage.ImportStubs.Count}, symbols={moduleImage.RuntimeSymbols.Count}");
}
catch (Exception ex)
{
failedModules++;
Log.Error($"Module load failed: {modulePath} ({ex.GetType().Name}: {ex.Message})", ex);
Console.Error.WriteLine($"[RUNTIME] Module load failed: {modulePath} ({ex.GetType().Name}: {ex.Message})");
}
}
Log.Info(
$"Module preload summary: loaded={loadedModules}, failed={failedModules}, merged_imports={mergedImportCount}, merged_symbols={mergedSymbolCount}");
Console.Error.WriteLine(
$"[RUNTIME] Module preload summary: loaded={loadedModules}, failed={failedModules}, merged_imports={mergedImportCount}, merged_symbols={mergedSymbolCount}");
return loadedImages;
}
private static void InstallNativePluginCompatibilityHooks(
IDictionary<ulong, string> importStubs,
SelfImage moduleImage,
string modulePath)
{
if (!string.Equals(Path.GetFileName(modulePath), "libfmod.prx", StringComparison.OrdinalIgnoreCase))
{
return;
}
ReadOnlySpan<string> compatibilityNids = ["uPLTdl3psGk"];
foreach (var nid in compatibilityNids)
{
if (moduleImage.RuntimeSymbols.TryGetValue(nid, out var address) && address >= 0x10000)
{
importStubs[address] = nid;
Console.Error.WriteLine(
$"[RUNTIME] Installed FMOD compatibility hook: {nid} -> 0x{address:X16}");
}
}
}
private void RebindImportedDataSymbols(
SelfImage mainImage,
IReadOnlyList<LoadedModuleImage> loadedModuleImages,
@@ -716,8 +781,8 @@ public sealed class SharpEmuRuntime : ISharpEmuRuntime
if (rebound != 0 || unresolved != 0)
{
Log.Info(
$"Imported data rebind: rebound={rebound}, unresolved={unresolved}");
Console.Error.WriteLine(
$"[RUNTIME] Imported data rebind: rebound={rebound}, unresolved={unresolved}");
}
}
@@ -749,8 +814,8 @@ public sealed class SharpEmuRuntime : ISharpEmuRuntime
{
if (logRebind)
{
Log.Warning(
$"Imported data unresolved: nid={relocation.Nid} target=0x{relocation.TargetAddress:X16} addend=0x{unchecked((ulong)relocation.Addend):X16}");
Console.Error.WriteLine(
$"[RUNTIME] Imported data unresolved: nid={relocation.Nid} target=0x{relocation.TargetAddress:X16} addend=0x{unchecked((ulong)relocation.Addend):X16}");
}
unresolved++;
@@ -762,8 +827,8 @@ public sealed class SharpEmuRuntime : ISharpEmuRuntime
{
if (logRebind)
{
Log.Error(
$"Imported data write-failed: nid={relocation.Nid} target=0x{relocation.TargetAddress:X16} value=0x{reboundValue:X16}");
Console.Error.WriteLine(
$"[RUNTIME] Imported data write-failed: nid={relocation.Nid} target=0x{relocation.TargetAddress:X16} value=0x{reboundValue:X16}");
}
unresolved++;
@@ -772,8 +837,8 @@ public sealed class SharpEmuRuntime : ISharpEmuRuntime
if (logRebind)
{
Log.Debug(
$"Imported data rebound: nid={relocation.Nid} target=0x{relocation.TargetAddress:X16} value=0x{reboundValue:X16}");
Console.Error.WriteLine(
$"[RUNTIME] Imported data rebound: nid={relocation.Nid} target=0x{relocation.TargetAddress:X16} value=0x{reboundValue:X16}");
}
rebound++;
@@ -809,8 +874,8 @@ public sealed class SharpEmuRuntime : ISharpEmuRuntime
{
if (!string.Equals(existingNid, nid, StringComparison.Ordinal))
{
Log.Warning(
$"Import stub conflict at 0x{address:X16}: keep={existingNid}, skip={nid} ({Path.GetFileName(modulePath)})");
Console.Error.WriteLine(
$"[RUNTIME] Import stub conflict at 0x{address:X16}: keep={existingNid}, skip={nid} ({Path.GetFileName(modulePath)})");
}
continue;
@@ -904,7 +969,7 @@ public sealed class SharpEmuRuntime : ISharpEmuRuntime
return !PreloadSkipModules.Contains(fileName);
}
private static void RegisterLoadedModule(string modulePath, SelfImage image, bool isMain, bool isSystemModule)
private int RegisterLoadedModule(string modulePath, SelfImage image, bool isMain, bool isSystemModule)
{
if (!TryComputeImageRange(image, out var baseAddress, out var size))
{
@@ -912,15 +977,27 @@ public sealed class SharpEmuRuntime : ISharpEmuRuntime
size = 0;
}
_ = TryGetEhFrameInfo(
image,
size,
out var ehFrameHeaderAddress,
out var ehFrameAddress,
out var ehFrameSize);
var handle = KernelModuleRegistry.RegisterModule(
modulePath,
baseAddress,
size,
image.EntryPoint,
image.InitFunctionEntryPoint,
ehFrameHeaderAddress,
ehFrameAddress,
ehFrameSize,
isMain,
isSystemModule);
Log.Info(
$"Registered module handle={handle} name={Path.GetFileName(modulePath)} base=0x{baseAddress:X16} size=0x{size:X16}");
KernelModuleRegistry.RegisterModuleSymbols(handle, image.RuntimeSymbols);
Console.Error.WriteLine(
$"[RUNTIME] Registered module handle={handle} name={Path.GetFileName(modulePath)} base=0x{baseAddress:X16} size=0x{size:X16}");
return handle;
}
private static bool TryComputeImageRange(SelfImage image, out ulong baseAddress, out ulong size)
@@ -1087,6 +1164,16 @@ public sealed class SharpEmuRuntime : ISharpEmuRuntime
disposableDispatcher.Dispose();
}
if (_cpuDispatcher is CpuDispatcher { NativeSessionLeaked: true })
{
// A guest worker is still inside guest code; unmapping the guest
// address space under it would fault the whole process, which
// hosts the GUI launcher in embedded sessions.
Console.Error.WriteLine(
"[RUNTIME] Guest workers were still active at teardown; keeping the guest address space mapped.");
return;
}
if (_virtualMemory is IDisposable disposableMemory)
{
disposableMemory.Dispose();
@@ -4,6 +4,7 @@
namespace SharpEmu.Core.Runtime;
using SharpEmu.Core.Cpu;
using SharpEmu.Core.Cpu.Debugging;
public readonly struct SharpEmuRuntimeOptions
{
@@ -12,4 +13,11 @@ public readonly struct SharpEmuRuntimeOptions
public bool StrictDynlibResolution { get; init; }
public int ImportTraceLimit { get; init; }
/// <summary>
/// An optional debugger to attach to guest execution. Flows through to
/// <see cref="CpuExecutionOptions.DebugHook"/>. Null (the default) runs with
/// no debugger attached.
/// </summary>
public ICpuDebugHook? DebugHook { get; init; }
}
+4
View File
@@ -14,6 +14,10 @@ SPDX-License-Identifier: GPL-2.0-or-later
<PackageReference Include="Iced" />
</ItemGroup>
<ItemGroup>
<InternalsVisibleTo Include="SharpEmu.Libs.Tests" />
</ItemGroup>
<PropertyGroup>
<NoWarn>$(NoWarn);1591</NoWarn>
</PropertyGroup>
-127
View File
@@ -1,127 +0,0 @@
{
"version": 2,
"dependencies": {
"net10.0": {
"Iced": {
"type": "Direct",
"requested": "[1.21.0, )",
"resolved": "1.21.0",
"contentHash": "dv5+81Q1TBQvVMSOOOmRcjJmvWcX3BZPZsIq31+RLc5cNft0IHAyNlkdb7ZarOWG913PyBoFDsDXoCIlKmLclg=="
},
"Microsoft.DotNet.PlatformAbstractions": {
"type": "Transitive",
"resolved": "3.1.6",
"contentHash": "jek4XYaQ/PGUwDKKhwR8K47Uh1189PFzMeLqO83mXrXQVIpARZCcfuDedH50YDTepBkfijCZN5U/vZi++erxtg=="
},
"Microsoft.Extensions.DependencyModel": {
"type": "Transitive",
"resolved": "9.0.9",
"contentHash": "fNGvKct2De8ghm0Bpfq0iWthtzIWabgOTi+gJhNOPhNJIowXNEUE2eZNW/zNCzrHVA3PXg2yZ+3cWZndC2IqYA=="
},
"Silk.NET.Core": {
"type": "Transitive",
"resolved": "2.23.0",
"contentHash": "D7AT/nnwlB+4RZ84XY8QNGBZMJI5z9l4CSSETIJ1wCfRJzRt/341y3MRZ4HbnFz4r/IGaWOEZr86iE+0/65yyQ==",
"dependencies": {
"Microsoft.DotNet.PlatformAbstractions": "3.1.6",
"Microsoft.Extensions.DependencyModel": "9.0.9"
}
},
"Silk.NET.GLFW": {
"type": "Transitive",
"resolved": "2.23.0",
"contentHash": "UIs4sH57xlPUNHQ/1bt9rymPWlGy8IMDCNv86h0iM4TOA1CkIx0XM/n/tA4AReh1zQkNrvkxPEdZ3Blvy1dyXg==",
"dependencies": {
"Silk.NET.Core": "2.23.0",
"Ultz.Native.GLFW": "3.4.0"
}
},
"Silk.NET.Maths": {
"type": "Transitive",
"resolved": "2.23.0",
"contentHash": "r8PdIVzME8EH0qAgbmRPO87I4GfgR2j8TofT7EMuRJDf1QluoQwnVypDoFJjQ2ZBSRsGYk5unYxxogI05Ogsmw=="
},
"Silk.NET.Windowing.Common": {
"type": "Transitive",
"resolved": "2.23.0",
"contentHash": "ThStSinmY9KQI8DGiF5XEhkLJVnBcgRTBTzL9ijg1wMZAYuckz7ykrNw04fjRm2Gryh6tCNGbvz2XaY0efeFzg==",
"dependencies": {
"Silk.NET.Core": "2.23.0",
"Silk.NET.Maths": "2.23.0"
}
},
"Silk.NET.Windowing.Glfw": {
"type": "Transitive",
"resolved": "2.23.0",
"contentHash": "aYBudKmENmvLRn9p15HbdvlQTnnXskcDfTfbYwSb/4fr263rGLwYuDw/txUEc2jihHJiWCp5+75Y7z5wTJWl7g==",
"dependencies": {
"Silk.NET.GLFW": "2.23.0",
"Silk.NET.Windowing.Common": "2.23.0"
}
},
"Ultz.Native.GLFW": {
"type": "Transitive",
"resolved": "3.4.0",
"contentHash": "Iy22JopynbOJ32vA0lBhFEzGi65GQJBuJHYBYRBpydrDpNoTiHnjIXfA65Gu+8qsOr/ZEoIF8r9aHCgAXuO6DA=="
},
"sharpemu.hle": {
"type": "Project",
"dependencies": {
"SharpEmu.Logging": "[1.0.0, )"
}
},
"sharpemu.libs": {
"type": "Project",
"dependencies": {
"SharpEmu.HLE": "[1.0.0, )",
"Silk.NET.Vulkan": "[2.23.0, )",
"Silk.NET.Vulkan.Extensions.EXT": "[2.23.0, )",
"Silk.NET.Vulkan.Extensions.KHR": "[2.23.0, )",
"Silk.NET.Windowing": "[2.23.0, )"
}
},
"sharpemu.logging": {
"type": "Project"
},
"Silk.NET.Vulkan": {
"type": "CentralTransitive",
"requested": "[2.23.0, )",
"resolved": "2.23.0",
"contentHash": "3/irtlSWXZ3eTi8N6nelI6L34NTB8ZJHpqVMNzZx2aX7Ek9YEQ34NoQW8/Tljrtmkg8KRhHW8hKTEzZaKV8PgA==",
"dependencies": {
"Silk.NET.Core": "2.23.0"
}
},
"Silk.NET.Vulkan.Extensions.EXT": {
"type": "CentralTransitive",
"requested": "[2.23.0, )",
"resolved": "2.23.0",
"contentHash": "+Oth189ksRiL6HvGCwIdnsYHawqrbO8y49u1H61z3wsfcHhQZeVDYe/wF5LD7fk3NcdgDvwFD3mLm1QWhdZySw==",
"dependencies": {
"Silk.NET.Core": "2.23.0",
"Silk.NET.Vulkan": "2.23.0"
}
},
"Silk.NET.Vulkan.Extensions.KHR": {
"type": "CentralTransitive",
"requested": "[2.23.0, )",
"resolved": "2.23.0",
"contentHash": "uRaf4j+SmH3DumjSSSUbFg33BnsGZUyXGj93O9NgGKZSJN3OTmNmQDxRew+/KiVLcgH6qzbto8aNGZ++j9GFWg==",
"dependencies": {
"Silk.NET.Core": "2.23.0",
"Silk.NET.Vulkan": "2.23.0"
}
},
"Silk.NET.Windowing": {
"type": "CentralTransitive",
"requested": "[2.23.0, )",
"resolved": "2.23.0",
"contentHash": "OPNPmt/lRyUKVYrFLQXVxyATqD3MKLc1iY1oKx1/2GppgmZxVZPwN12tekrQ4C7408kgB1L5JD1Wnirqqeb2kg==",
"dependencies": {
"Silk.NET.Windowing.Common": "2.23.0",
"Silk.NET.Windowing.Glfw": "2.23.0"
}
}
}
}
}
@@ -0,0 +1,54 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using System.Net;
namespace SharpEmu.DebugClient;
/// <summary>
/// Parses the <c>host:port</c> the client connects to. Mirrors the server's
/// defaults (loopback, port 5714) so a bare invocation attaches to a local
/// emulator with no arguments.
/// </summary>
internal static class ClientEndpoint
{
public const int DefaultPort = 5714;
public static bool TryParse(string? text, out string host, out int port, out string error)
{
host = "127.0.0.1";
port = DefaultPort;
error = string.Empty;
if (string.IsNullOrWhiteSpace(text))
{
return true;
}
var value = text.Trim();
var separator = value.LastIndexOf(':');
if (separator >= 0)
{
var portText = value[(separator + 1)..];
if (portText.Length > 0 && (!int.TryParse(portText, out port) || port is <= 0 or > 65535))
{
error = $"Invalid port '{portText}'.";
return false;
}
value = value[..separator];
}
if (!string.IsNullOrWhiteSpace(value))
{
host = string.Equals(value, "localhost", StringComparison.OrdinalIgnoreCase) ? "127.0.0.1" : value;
}
if (!IPAddress.TryParse(host, out _) && !Uri.CheckHostName(host).Equals(UriHostNameType.Dns))
{
error = $"Invalid host '{host}'.";
return false;
}
return true;
}
}
@@ -0,0 +1,160 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using System.Text.Json;
namespace SharpEmu.DebugClient;
/// <summary>
/// Turns a friendly REPL line (<c>mem 0x1000 64</c>) into the JSON request the
/// server understands. Local-only verbs (help, quit) are reported back to the
/// caller instead of producing a request.
/// </summary>
internal static class CommandTranslator
{
public enum ActionKind
{
SendRequest,
ShowHelp,
Quit,
Ignore,
Error,
}
public readonly record struct Result(ActionKind Kind, string? Payload = null, string? Error = null);
public static Result Translate(string line)
{
var trimmed = line.Trim();
if (trimmed.Length == 0)
{
return new Result(ActionKind.Ignore);
}
var parts = trimmed.Split((char[]?)null, StringSplitOptions.RemoveEmptyEntries);
var verb = parts[0].ToLowerInvariant();
switch (verb)
{
case "help" or "?":
return new Result(ActionKind.ShowHelp);
case "quit" or "exit" or "q":
return new Result(ActionKind.Quit);
case "raw":
var json = trimmed[verb.Length..].Trim();
return json.Length == 0
? Error("raw requires a JSON object argument.")
: Send(json);
case "ping":
return Request("ping");
case "status" or "info":
return Request("status");
case "state":
return Request("state");
case "regs" or "registers":
return Request("registers");
case "continue" or "cont" or "c":
return Request("continue");
case "step" or "s":
return Request("step");
case "pause" or "p":
return Request("pause");
case "bp" or "breakpoints" or "bl":
return Request("list-breakpoints");
case "setreg":
return parts.Length >= 3
? Request("set-register", ("register", parts[1]), ("value", parts[2]))
: Error("Usage: setreg <register> <value>");
case "mem" or "read":
return parts.Length >= 3
? Request("read-memory", ("address", parts[1]), ("length", parts[2]))
: Error("Usage: mem <address> <length>");
case "write":
return parts.Length >= 3
? Request("write-memory", ("address", parts[1]), ("bytes", parts[2]))
: Error("Usage: write <address> <hex-bytes>");
case "break" or "b":
if (parts.Length < 2)
{
return Error("Usage: break <address> [kind] [length]");
}
var breakArgs = new List<(string, string)> { ("address", parts[1]) };
if (parts.Length >= 3)
{
breakArgs.Add(("kind", parts[2]));
}
if (parts.Length >= 4)
{
breakArgs.Add(("length", parts[3]));
}
return Request("add-breakpoint", breakArgs.ToArray());
case "del" or "rm" or "delete":
return parts.Length >= 2
? Request("remove-breakpoint", ("id", parts[1]))
: Error("Usage: del <id>");
case "enable":
return parts.Length >= 2
? Request("enable-breakpoint", ("id", parts[1]), ("enabled", "true"))
: Error("Usage: enable <id>");
case "disable":
return parts.Length >= 2
? RequestWithBool("enable-breakpoint", ("id", parts[1]), enabledName: "enabled", enabled: false)
: Error("Usage: disable <id>");
default:
return Error($"Unknown command '{verb}'. Type 'help' for the command list.");
}
}
private static Result Request(string command, params (string Name, string Value)[] args)
{
var payload = new Dictionary<string, object?> { ["command"] = command };
foreach (var (name, value) in args)
{
payload[name] = value;
}
return Send(JsonSerializer.Serialize(payload));
}
private static Result RequestWithBool(string command, (string Name, string Value) idArg, string enabledName, bool enabled)
{
var payload = new Dictionary<string, object?>
{
["command"] = command,
[idArg.Name] = idArg.Value,
[enabledName] = enabled,
};
return Send(JsonSerializer.Serialize(payload));
}
private static Result Send(string json) => new(ActionKind.SendRequest, json);
private static Result Error(string message) => new(ActionKind.Error, Error: message);
public const string HelpText = """
SharpEmu debug client commands:
status | info Show target state and last stop
state Show run state only
regs | registers Dump integer registers (paused only)
setreg <reg> <value> Set a register (rip/rflags/gp, paused only)
mem <addr> <len> Read guest memory as hex (paused only)
write <addr> <hex> Write guest memory from hex (paused only)
break <addr> [kind] [len] Add a breakpoint (kind: execute/readwatch/writewatch/accesswatch)
bp | breakpoints List breakpoints
del <id> Remove a breakpoint
enable <id> / disable <id> Toggle a breakpoint
continue | c Resume the target
step | s Resume and stop at the next frame
pause Ask a running target to stop
ping Round-trip check
raw <json> Send a literal JSON request
help | ? Show this help
quit | exit Disconnect and exit
Addresses and values accept decimal or 0x-prefixed hex.
""";
}
+153
View File
@@ -0,0 +1,153 @@
<!--
Copyright (C) 2026 SharpEmu Emulator Project
SPDX-License-Identifier: GPL-2.0-or-later
-->
# SharpEmu.DebugClient
A small, standalone command-line client that connects to the SharpEmu
emulator's **live debug server** and drives it interactively. It ships as its
own executable (`SharpEmu.DebugClient`) and takes no dependency on the emulator
assemblies — it speaks the server's line-delimited JSON protocol directly over
TCP, so you can also drive the server from `nc`, a script, or your own tool.
> **Status:** infrastructure. The transport, protocol, session model, and
> breakpoint store are in place. Stops are delivered at **frame boundaries**
> (process entry and each module initializer); per-instruction stepping and data
> watchpoints are part of the surface and become live as the CPU backend grows
> the corresponding hooks. See [`docs/debugger-server.md`](../../docs/debugger-server.md)
> for the architecture and protocol reference.
## How it fits together
```
+-------------------------+ TCP (JSON lines) +----------------------+
| SharpEmu (emulator) | <------------------------------> | SharpEmu.DebugClient |
| --debug-server | | (this executable) |
| | | |
| DebuggerServerHost | | REPL / --exec |
| +- DebuggerServer | frame boundaries via ICpuDebugHook| |
| +- DebuggerSession <-+------ CPU dispatcher ------------ | |
+-------------------------+ +----------------------+
```
The emulator is the **server**; this client is a separate process that connects
to it and issues commands. The two never share memory — everything crosses the
socket as JSON.
## Building
```bash
dotnet build src/SharpEmu.DebugClient/SharpEmu.DebugClient.csproj
```
## Quick start
1. Launch the emulator with the debug server enabled. It listens on
`127.0.0.1:5714` by default and, with stop-at-entry on, parks the guest at
its first frame until you continue:
```bash
SharpEmu --debug-server "/path/to/game/eboot.bin"
# or choose an endpoint:
SharpEmu --debug-server=127.0.0.1:5714 "/path/to/game/eboot.bin"
```
2. In another terminal, attach the client:
```bash
SharpEmu.DebugClient # defaults to 127.0.0.1:5714
SharpEmu.DebugClient 127.0.0.1:5714 # explicit endpoint
```
3. Drive the target:
```
status
regs
break 0x00000008801234a0
continue
mem 0x00000008802000000 64
```
## Invocation
```
SharpEmu.DebugClient [host:port] [--exec "<command>"]... [--quiet]
```
| Option | Meaning |
| ------------- | ------------------------------------------------------------- |
| `host:port` | Server endpoint. Default `127.0.0.1:5714`. `localhost` is fine. |
| `--exec, -e` | Run one command non-interactively, then exit. Repeatable. |
| `--quiet` | Suppress the connection banner. |
| `--help, -h` | Show usage and the command list. |
Non-interactive example (scriptable):
```bash
SharpEmu.DebugClient --exec "break 0x8801234a0" --exec "continue"
```
## Commands
Addresses and values accept decimal or `0x`-prefixed hex. Register and memory
commands only succeed while the target is **paused**.
| Command | Server verb | Description |
| ------- | ----------- | ----------- |
| `status` \| `info` | `status` | Target state plus the last stop. |
| `state` | `state` | Run state only (`Running`/`Paused`/…). |
| `regs` \| `registers` | `registers` | Dump the integer registers. |
| `setreg <reg> <value>` | `set-register` | Set `rip`, `rflags`, or a GP register. |
| `mem <addr> <len>` \| `read <addr> <len>` | `read-memory` | Read guest memory as hex. |
| `write <addr> <hex>` | `write-memory` | Write guest memory from a hex string. |
| `break <addr> [kind] [len]` \| `b …` | `add-breakpoint` | Add a breakpoint. `kind`: `execute` (default), `readwatch`, `writewatch`, `accesswatch`. |
| `bp` \| `breakpoints` | `list-breakpoints` | List breakpoints. |
| `del <id>` \| `rm <id>` | `remove-breakpoint` | Remove a breakpoint. |
| `enable <id>` / `disable <id>` | `enable-breakpoint` | Toggle a breakpoint. |
| `continue` \| `c` | `continue` | Resume a paused target. |
| `step` \| `s` | `step` | Resume and stop at the next frame boundary. |
| `pause` | `pause` | Ask a running target to stop at the next boundary. |
| `ping` | `ping` | Round-trip liveness check. |
| `raw <json>` | *(passthrough)* | Send a literal JSON request. |
| `help` \| `?` | — | Show the command list (local). |
| `quit` \| `exit` | — | Disconnect and exit (local). |
## Output
The client prints two kinds of lines as they arrive:
- `reply>` — the response to a command you sent (`ok`, plus `data` or `error`).
- `event>` — an unsolicited notification: `hello` on connect, `stopped` when the
target hits a breakpoint / entry / step / pause, `resumed` on continue, and
`terminated` when the run ends.
Because replies and events share one stream, the client prints everything it
receives rather than pairing replies to requests — a `stopped` event may arrive
between your command and its reply.
## Protocol (for building your own client)
One JSON object per line, UTF-8, `\n`-terminated, in both directions.
Request:
```json
{"command":"read-memory","address":"0x8802000000","length":64}
```
Reply:
```json
{"ok":true,"command":"read-memory","data":{"address":"0x0000000880200000","length":64,"bytes":"48894C24.."}}
```
Event:
```json
{"event":"stopped","reason":"Breakpoint","address":"0x00000008801234A0","frameKind":"ProcessEntry","frameLabel":"eboot.bin","registers":{ ... }}
```
The full verb list and payload fields live in
[`docs/debugger-server.md`](../../docs/debugger-server.md).
@@ -0,0 +1,120 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using System.Net.Sockets;
using System.Text;
using System.Text.Json;
namespace SharpEmu.DebugClient;
/// <summary>
/// A thin TCP wrapper around the server's line-delimited JSON protocol: it
/// writes request lines and runs a background loop that prints incoming
/// responses and events as they arrive. Because the stream interleaves replies
/// with asynchronous stop/resume events, a single reader printing everything is
/// simpler and more robust than correlating request/response pairs.
/// </summary>
internal sealed class DebugClientConnection : IAsyncDisposable
{
private static readonly UTF8Encoding Utf8NoBom = new(encoderShouldEmitUTF8Identifier: false);
private static readonly JsonSerializerOptions PrettyOptions = new() { WriteIndented = true };
private readonly TcpClient _client;
private readonly StreamReader _reader;
private readonly StreamWriter _writer;
private readonly SemaphoreSlim _writeLock = new(1, 1);
private DebugClientConnection(TcpClient client, NetworkStream stream)
{
_client = client;
_reader = new StreamReader(stream, Utf8NoBom);
_writer = new StreamWriter(stream, Utf8NoBom) { AutoFlush = false };
}
public static async Task<DebugClientConnection> ConnectAsync(string host, int port, CancellationToken cancellationToken)
{
var client = new TcpClient();
await client.ConnectAsync(host, port, cancellationToken).ConfigureAwait(false);
return new DebugClientConnection(client, client.GetStream());
}
/// <summary>Continuously prints incoming lines until the stream closes.</summary>
public async Task ReceiveLoopAsync(CancellationToken cancellationToken)
{
try
{
while (!cancellationToken.IsCancellationRequested)
{
var line = await _reader.ReadLineAsync(cancellationToken).ConfigureAwait(false);
if (line is null)
{
Console.WriteLine();
Console.WriteLine("[connection closed by server]");
return;
}
Print(line);
}
}
catch (OperationCanceledException)
{
}
catch (IOException)
{
Console.WriteLine();
Console.WriteLine("[connection lost]");
}
}
public async Task SendAsync(string json, CancellationToken cancellationToken)
{
await _writeLock.WaitAsync(cancellationToken).ConfigureAwait(false);
try
{
await _writer.WriteLineAsync(json.AsMemory(), cancellationToken).ConfigureAwait(false);
await _writer.FlushAsync(cancellationToken).ConfigureAwait(false);
}
finally
{
_writeLock.Release();
}
}
private static void Print(string line)
{
try
{
using var document = JsonDocument.Parse(line);
var root = document.RootElement;
var isEvent = root.TryGetProperty("event", out _);
var prefix = isEvent ? "event>" : "reply>";
var pretty = JsonSerializer.Serialize(root, PrettyOptions);
Console.WriteLine();
Console.WriteLine($"{prefix}\n{pretty}");
}
catch (JsonException)
{
Console.WriteLine();
Console.WriteLine(line);
}
}
public async ValueTask DisposeAsync()
{
try
{
await _writer.FlushAsync().ConfigureAwait(false);
}
catch (IOException)
{
}
catch (ObjectDisposedException)
{
}
_writeLock.Dispose();
_reader.Dispose();
await _writer.DisposeAsync().ConfigureAwait(false);
_client.Dispose();
}
}
+193
View File
@@ -0,0 +1,193 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using System.Net.Sockets;
using SharpEmu.DebugClient;
return await ClientProgram.RunAsync(args).ConfigureAwait(false);
internal static class ClientProgram
{
public static async Task<int> RunAsync(string[] args)
{
if (args.Any(a => a is "--help" or "-h"))
{
PrintUsage();
return 0;
}
string? endpointArg = null;
var execCommands = new List<string>();
var quiet = false;
for (var i = 0; i < args.Length; i++)
{
var arg = args[i];
if (string.Equals(arg, "--exec", StringComparison.OrdinalIgnoreCase) || string.Equals(arg, "-e", StringComparison.OrdinalIgnoreCase))
{
if (i + 1 >= args.Length)
{
Console.Error.WriteLine("--exec requires a command argument.");
return 2;
}
execCommands.Add(args[++i]);
continue;
}
if (string.Equals(arg, "--quiet", StringComparison.OrdinalIgnoreCase))
{
quiet = true;
continue;
}
if (arg.StartsWith('-'))
{
Console.Error.WriteLine($"Unknown option '{arg}'.");
PrintUsage();
return 2;
}
endpointArg ??= arg;
}
if (!ClientEndpoint.TryParse(endpointArg, out var host, out var port, out var endpointError))
{
Console.Error.WriteLine(endpointError);
return 2;
}
using var shutdown = new CancellationTokenSource();
Console.CancelKeyPress += (_, eventArgs) =>
{
eventArgs.Cancel = true;
shutdown.Cancel();
};
DebugClientConnection connection;
try
{
connection = await DebugClientConnection.ConnectAsync(host, port, shutdown.Token).ConfigureAwait(false);
}
catch (Exception ex) when (ex is SocketException or OperationCanceledException)
{
Console.Error.WriteLine($"Could not connect to {host}:{port}: {ex.Message}");
Console.Error.WriteLine("Start the emulator with --debug-server first.");
return 3;
}
await using (connection)
{
var receiveTask = connection.ReceiveLoopAsync(shutdown.Token);
if (execCommands.Count > 0)
{
await RunOneShotAsync(connection, execCommands, shutdown.Token).ConfigureAwait(false);
}
else
{
if (!quiet)
{
Console.WriteLine($"Connected to SharpEmu debug server at {host}:{port}.");
Console.WriteLine("Type 'help' for commands, 'quit' to exit.");
}
await RunReplAsync(connection, shutdown).ConfigureAwait(false);
}
shutdown.Cancel();
try
{
await receiveTask.ConfigureAwait(false);
}
catch (OperationCanceledException)
{
}
}
return 0;
}
private static async Task RunOneShotAsync(
DebugClientConnection connection,
IReadOnlyList<string> commands,
CancellationToken cancellationToken)
{
foreach (var command in commands)
{
var result = CommandTranslator.Translate(command);
switch (result.Kind)
{
case CommandTranslator.ActionKind.SendRequest:
await connection.SendAsync(result.Payload!, cancellationToken).ConfigureAwait(false);
break;
case CommandTranslator.ActionKind.Error:
Console.Error.WriteLine(result.Error);
break;
case CommandTranslator.ActionKind.ShowHelp:
Console.WriteLine(CommandTranslator.HelpText);
break;
}
}
// Give the server a moment to answer before the client exits.
try
{
await Task.Delay(TimeSpan.FromMilliseconds(500), cancellationToken).ConfigureAwait(false);
}
catch (OperationCanceledException)
{
}
}
private static async Task RunReplAsync(DebugClientConnection connection, CancellationTokenSource shutdown)
{
while (!shutdown.IsCancellationRequested)
{
var line = await Console.In.ReadLineAsync(shutdown.Token).ConfigureAwait(false);
if (line is null)
{
break;
}
var result = CommandTranslator.Translate(line);
switch (result.Kind)
{
case CommandTranslator.ActionKind.Quit:
return;
case CommandTranslator.ActionKind.ShowHelp:
Console.WriteLine(CommandTranslator.HelpText);
break;
case CommandTranslator.ActionKind.Error:
Console.Error.WriteLine(result.Error);
break;
case CommandTranslator.ActionKind.Ignore:
break;
case CommandTranslator.ActionKind.SendRequest:
try
{
await connection.SendAsync(result.Payload!, shutdown.Token).ConfigureAwait(false);
}
catch (Exception ex) when (ex is IOException or ObjectDisposedException)
{
Console.Error.WriteLine("Send failed; the connection is closed.");
return;
}
break;
}
}
}
private static void PrintUsage()
{
Console.WriteLine("SharpEmu.DebugClient — live debugger client for the SharpEmu debug server.");
Console.WriteLine();
Console.WriteLine("Usage: SharpEmu.DebugClient [host:port] [--exec \"<command>\"]... [--quiet]");
Console.WriteLine(" host:port Server endpoint (default 127.0.0.1:5714).");
Console.WriteLine(" --exec, -e Run a command non-interactively (repeatable), then exit.");
Console.WriteLine(" --quiet Suppress the connection banner.");
Console.WriteLine(" --help, -h Show this help.");
Console.WriteLine();
Console.WriteLine(CommandTranslator.HelpText);
}
}
@@ -0,0 +1,16 @@
<!--
Copyright (C) 2026 SharpEmu Emulator Project
SPDX-License-Identifier: GPL-2.0-or-later
-->
<Project Sdk="Microsoft.NET.Sdk">
<PropertyGroup>
<!-- A small, standalone console tool: it speaks the debug server's
line-delimited JSON protocol directly over TCP and takes no dependency
on the emulator assemblies, so it builds and ships independently. -->
<OutputType>Exe</OutputType>
<AssemblyName>SharpEmu.DebugClient</AssemblyName>
<RootNamespace>SharpEmu.DebugClient</RootNamespace>
<NoWarn>$(NoWarn);1591</NoWarn>
</PropertyGroup>
</Project>
@@ -0,0 +1,48 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
namespace SharpEmu.Debugger.Breakpoints;
/// <summary>
/// A single breakpoint or watchpoint. Instances are immutable; the owning
/// <see cref="BreakpointStore"/> replaces an entry to change its enabled state.
/// </summary>
public sealed class Breakpoint
{
public Breakpoint(int id, BreakpointKind kind, ulong address, ulong length = 1, bool enabled = true)
{
if (length == 0)
{
throw new ArgumentOutOfRangeException(nameof(length), "Breakpoint length must be at least one byte.");
}
Id = id;
Kind = kind;
Address = address;
Length = length;
Enabled = enabled;
}
/// <summary>The store-assigned identifier used by clients to reference it.</summary>
public int Id { get; }
public BreakpointKind Kind { get; }
/// <summary>The first guest address the breakpoint covers.</summary>
public ulong Address { get; }
/// <summary>
/// The number of bytes the breakpoint covers. Always one for
/// <see cref="BreakpointKind.Execute"/>; the watch kinds may span a range.
/// </summary>
public ulong Length { get; }
public bool Enabled { get; }
/// <summary>True when <paramref name="address"/> falls within this breakpoint.</summary>
public bool Covers(ulong address) => address >= Address && address < Address + Length;
/// <summary>Returns a copy with a different enabled state.</summary>
public Breakpoint WithEnabled(bool enabled)
=> enabled == Enabled ? this : new Breakpoint(Id, Kind, Address, Length, enabled);
}
@@ -0,0 +1,26 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
namespace SharpEmu.Debugger.Breakpoints;
/// <summary>
/// The kind of stop a breakpoint requests. Execution breakpoints are honoured
/// at the frame-boundary seam that exists today; the data-watch kinds are part
/// of the surface so client protocols and tooling can be built against them,
/// and are armed once the execution backend can report the corresponding
/// accesses.
/// </summary>
public enum BreakpointKind
{
/// <summary>Stop when the instruction pointer reaches the address.</summary>
Execute,
/// <summary>Stop when the guest reads from the address range.</summary>
ReadWatch,
/// <summary>Stop when the guest writes to the address range.</summary>
WriteWatch,
/// <summary>Stop when the guest reads from or writes to the address range.</summary>
AccessWatch,
}
@@ -0,0 +1,90 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
namespace SharpEmu.Debugger.Breakpoints;
/// <summary>
/// A thread-safe registry of breakpoints. The debug server mutates it from
/// client-servicing threads while the emulation thread queries it at frame
/// boundaries, so every operation takes the same lock.
/// </summary>
public sealed class BreakpointStore
{
private readonly object _sync = new();
private readonly Dictionary<int, Breakpoint> _breakpoints = new();
private int _nextId = 1;
/// <summary>Adds a breakpoint and returns the created entry with its id.</summary>
public Breakpoint Add(BreakpointKind kind, ulong address, ulong length = 1)
{
lock (_sync)
{
var effectiveLength = kind == BreakpointKind.Execute ? 1UL : Math.Max(1UL, length);
var breakpoint = new Breakpoint(_nextId++, kind, address, effectiveLength);
_breakpoints[breakpoint.Id] = breakpoint;
return breakpoint;
}
}
/// <summary>Removes a breakpoint by id. Returns false when it did not exist.</summary>
public bool Remove(int id)
{
lock (_sync)
{
return _breakpoints.Remove(id);
}
}
/// <summary>Enables or disables a breakpoint by id.</summary>
public bool SetEnabled(int id, bool enabled)
{
lock (_sync)
{
if (!_breakpoints.TryGetValue(id, out var breakpoint))
{
return false;
}
_breakpoints[id] = breakpoint.WithEnabled(enabled);
return true;
}
}
/// <summary>Removes every breakpoint.</summary>
public void Clear()
{
lock (_sync)
{
_breakpoints.Clear();
}
}
/// <summary>Returns a point-in-time copy of all breakpoints.</summary>
public IReadOnlyList<Breakpoint> Snapshot()
{
lock (_sync)
{
return _breakpoints.Values.ToArray();
}
}
/// <summary>
/// Finds the first enabled execution breakpoint covering <paramref name="address"/>,
/// or null when none applies.
/// </summary>
public Breakpoint? FindExecuteHit(ulong address)
{
lock (_sync)
{
foreach (var breakpoint in _breakpoints.Values)
{
if (breakpoint.Enabled && breakpoint.Kind == BreakpointKind.Execute && breakpoint.Covers(address))
{
return breakpoint;
}
}
return null;
}
}
}
@@ -0,0 +1,73 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using SharpEmu.Core.Cpu.Debugging;
using SharpEmu.HLE;
namespace SharpEmu.Debugger;
/// <summary>
/// An immutable snapshot of the guest integer register state at a stop. XMM/YMM
/// state is intentionally omitted here and read on demand through the target to
/// keep the common register-dump path cheap.
/// </summary>
public readonly struct DebugRegisterFile
{
private readonly ulong[] _generalPurpose;
public DebugRegisterFile(
ulong[] generalPurpose,
ulong rip,
ulong rflags,
ulong fsBase,
ulong gsBase)
{
ArgumentNullException.ThrowIfNull(generalPurpose);
if (generalPurpose.Length != 16)
{
throw new ArgumentException("Expected 16 general-purpose registers.", nameof(generalPurpose));
}
_generalPurpose = generalPurpose;
Rip = rip;
Rflags = rflags;
FsBase = fsBase;
GsBase = gsBase;
}
public ulong Rip { get; }
public ulong Rflags { get; }
public ulong FsBase { get; }
public ulong GsBase { get; }
/// <summary>Reads a register by identifier.</summary>
public ulong this[DebugRegisterId id] => id switch
{
DebugRegisterId.Rip => Rip,
DebugRegisterId.Rflags => Rflags,
DebugRegisterId.FsBase => FsBase,
DebugRegisterId.GsBase => GsBase,
_ when id.IsGeneralPurpose() => _generalPurpose[(int)id],
_ => throw new ArgumentOutOfRangeException(nameof(id), id, null),
};
/// <summary>Reads a general-purpose register.</summary>
public ulong this[CpuRegister register] => _generalPurpose[(int)register];
/// <summary>Captures the integer register state of a live debug frame.</summary>
public static DebugRegisterFile Capture(ICpuDebugFrame frame)
{
ArgumentNullException.ThrowIfNull(frame);
var gpr = new ulong[16];
for (var i = 0; i < gpr.Length; i++)
{
gpr[i] = frame.GetRegister((CpuRegister)i);
}
return new DebugRegisterFile(gpr, frame.Rip, frame.Rflags, frame.FsBase, frame.GsBase);
}
}
+62
View File
@@ -0,0 +1,62 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using SharpEmu.HLE;
namespace SharpEmu.Debugger;
/// <summary>
/// The registers a debugger can name. The first sixteen values line up with
/// <see cref="CpuRegister"/> so a general-purpose register can be converted
/// between the two enums by casting; the remaining values cover the special
/// registers a debug frame exposes.
/// </summary>
public enum DebugRegisterId
{
Rax = 0,
Rcx = 1,
Rdx = 2,
Rbx = 3,
Rsp = 4,
Rbp = 5,
Rsi = 6,
Rdi = 7,
R8 = 8,
R9 = 9,
R10 = 10,
R11 = 11,
R12 = 12,
R13 = 13,
R14 = 14,
R15 = 15,
Rip = 16,
Rflags = 17,
FsBase = 18,
GsBase = 19,
}
/// <summary>Helpers for mapping between debug and CPU register identifiers.</summary>
public static class DebugRegisterIdExtensions
{
/// <summary>
/// True when the identifier names one of the sixteen general-purpose
/// registers and can be cast to <see cref="CpuRegister"/>.
/// </summary>
public static bool IsGeneralPurpose(this DebugRegisterId id)
=> id is >= DebugRegisterId.Rax and <= DebugRegisterId.R15;
/// <summary>
/// Converts a general-purpose identifier to its <see cref="CpuRegister"/>.
/// Throws when <paramref name="id"/> is a special register.
/// </summary>
public static CpuRegister ToCpuRegister(this DebugRegisterId id)
{
if (!id.IsGeneralPurpose())
{
throw new ArgumentOutOfRangeException(nameof(id), id, "Not a general-purpose register.");
}
return (CpuRegister)(int)id;
}
}
@@ -0,0 +1,59 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using System.Net;
using SharpEmu.Core.Cpu.Debugging;
using SharpEmu.Debugger.Server;
using SharpEmu.Debugger.Session;
namespace SharpEmu.Debugger;
/// <summary>
/// One-call wiring of the live debugger: it owns a <see cref="DebuggerSession"/>
/// and a <see cref="DebuggerServer"/>, exposes the <see cref="Hook"/> to attach
/// to <c>SharpEmuRuntimeOptions.DebugHook</c>, and starts/stops the network
/// front-end. A host constructs one, hands <see cref="Hook"/> to the runtime,
/// calls <see cref="Start"/>, and calls <see cref="NotifyRunCompleted"/> once the
/// runtime returns.
/// </summary>
public sealed class DebuggerServerHost : IAsyncDisposable
{
private readonly DebuggerSession _session;
private readonly DebuggerServer _server;
public DebuggerServerHost(
DebuggerServerOptions? serverOptions = null,
DebuggerSessionOptions? sessionOptions = null)
{
_session = new DebuggerSession(sessionOptions);
_server = new DebuggerServer(_session, serverOptions);
}
/// <summary>The session driving the target.</summary>
public IDebuggerSession Session => _session;
/// <summary>
/// The dispatcher hook to hand to the runtime so guest frames route through
/// the debugger.
/// </summary>
public ICpuDebugHook Hook => _session.Hook;
/// <summary>The endpoint the server bound to, or null before <see cref="Start"/>.</summary>
public IPEndPoint? Endpoint => _server.Endpoint;
/// <summary>Begins accepting debugger clients.</summary>
public void Start() => _server.Start();
/// <summary>
/// Releases a parked emulation thread and marks the target terminated. Call
/// after the runtime's run returns so any attached client is notified and the
/// guest thread is never left blocked in the debugger.
/// </summary>
public void NotifyRunCompleted() => _session.NotifyTerminated();
public async ValueTask DisposeAsync()
{
_session.NotifyTerminated();
await _server.DisposeAsync().ConfigureAwait(false);
}
}
@@ -0,0 +1,340 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using SharpEmu.Debugger.Breakpoints;
using SharpEmu.Debugger.Session;
using SharpEmu.HLE;
namespace SharpEmu.Debugger.Protocol;
/// <summary>
/// Translates parsed <see cref="DebugRequest"/> verbs into operations on an
/// <see cref="IDebuggerSession"/> and packages the outcome as a
/// <see cref="DebugResponse"/>. This is the single place command semantics live,
/// so it is shared by every connection and independent of the wire format.
/// </summary>
public sealed class DebugCommandDispatcher
{
private readonly IDebuggerSession _session;
public DebugCommandDispatcher(IDebuggerSession session)
{
_session = session ?? throw new ArgumentNullException(nameof(session));
}
public DebugResponse Dispatch(DebugRequest request)
{
return request.Command switch
{
JsonLineDebugProtocol.ParseErrorCommand => ParseError(request),
"ping" => DebugResponse.Success(request.Command),
"status" or "info" => Status(request),
"state" => DebugResponse.Success(request.Command, new Dictionary<string, object?>
{
["state"] = _session.State.ToString(),
}),
"registers" or "regs" => Registers(request),
"set-register" or "set-reg" => SetRegister(request),
"read-memory" or "read-mem" => ReadMemory(request),
"write-memory" or "write-mem" => WriteMemory(request),
"list-breakpoints" or "breakpoints" => ListBreakpoints(request),
"add-breakpoint" or "break" => AddBreakpoint(request),
"remove-breakpoint" or "delete-breakpoint" => RemoveBreakpoint(request),
"enable-breakpoint" => EnableBreakpoint(request),
"continue" or "cont" or "c" => Simple(request, _session.Continue(), "Target is not paused."),
"step" or "s" => Simple(request, _session.StepFrame(), "Target is not paused."),
"pause" => Pause(request),
_ => DebugResponse.Failure(request.Command, $"Unknown command '{request.Command}'."),
};
}
private static DebugResponse ParseError(DebugRequest request)
{
var message = request.TryGetString("message", out var text) ? text : "Malformed request.";
return DebugResponse.Failure(request.Command, message);
}
private DebugResponse Status(DebugRequest request)
{
var data = new Dictionary<string, object?>
{
["state"] = _session.State.ToString(),
["breakpoints"] = _session.Breakpoints.Snapshot().Count,
};
if (_session.LastStop is { } stop)
{
data["lastStop"] = DescribeStop(stop);
}
return DebugResponse.Success(request.Command, data);
}
private DebugResponse Registers(DebugRequest request)
{
if (!_session.TryGetRegisters(out var registers))
{
return NotPaused(request);
}
return DebugResponse.Success(request.Command, new Dictionary<string, object?>
{
["registers"] = DescribeRegisters(registers),
});
}
private DebugResponse SetRegister(DebugRequest request)
{
if (!request.TryGetString("register", out var name) || !TryParseRegister(name, out var id))
{
return DebugResponse.Failure(request.Command, "Expected a valid 'register' name.");
}
if (!request.TryGetUInt64("value", out var value))
{
return DebugResponse.Failure(request.Command, "Expected a 'value'.");
}
return _session.TrySetRegister(id, value)
? DebugResponse.Success(request.Command)
: DebugResponse.Failure(request.Command, "Register is not writable or target is not paused.");
}
private DebugResponse ReadMemory(DebugRequest request)
{
if (!request.TryGetUInt64("address", out var address))
{
return DebugResponse.Failure(request.Command, "Expected an 'address'.");
}
if (!request.TryGetInt32("length", out var length) || length <= 0 || length > MaxMemoryChunk)
{
return DebugResponse.Failure(request.Command, $"Expected a 'length' between 1 and {MaxMemoryChunk}.");
}
var buffer = new byte[length];
if (!_session.TryReadMemory(address, buffer))
{
return DebugResponse.Failure(request.Command, "Memory is unreadable or target is not paused.");
}
return DebugResponse.Success(request.Command, new Dictionary<string, object?>
{
["address"] = FormatAddress(address),
["length"] = length,
["bytes"] = Convert.ToHexString(buffer),
});
}
private DebugResponse WriteMemory(DebugRequest request)
{
if (!request.TryGetUInt64("address", out var address))
{
return DebugResponse.Failure(request.Command, "Expected an 'address'.");
}
if (!request.TryGetString("bytes", out var hex) || hex.Length == 0 || (hex.Length & 1) != 0)
{
return DebugResponse.Failure(request.Command, "Expected 'bytes' as an even-length hex string.");
}
byte[] data;
try
{
data = Convert.FromHexString(hex);
}
catch (FormatException)
{
return DebugResponse.Failure(request.Command, "'bytes' is not valid hex.");
}
if (data.Length > MaxMemoryChunk)
{
return DebugResponse.Failure(request.Command, $"Cannot write more than {MaxMemoryChunk} bytes at once.");
}
return _session.TryWriteMemory(address, data)
? DebugResponse.Success(request.Command, new Dictionary<string, object?> { ["written"] = data.Length })
: DebugResponse.Failure(request.Command, "Memory is unwritable or target is not paused.");
}
private DebugResponse ListBreakpoints(DebugRequest request)
{
var breakpoints = _session.Breakpoints.Snapshot()
.OrderBy(breakpoint => breakpoint.Id)
.Select(DescribeBreakpoint)
.ToArray();
return DebugResponse.Success(request.Command, new Dictionary<string, object?>
{
["breakpoints"] = breakpoints,
});
}
private DebugResponse AddBreakpoint(DebugRequest request)
{
if (!request.TryGetUInt64("address", out var address))
{
return DebugResponse.Failure(request.Command, "Expected an 'address'.");
}
var kind = BreakpointKind.Execute;
if (request.TryGetString("kind", out var kindText) && !TryParseBreakpointKind(kindText, out kind))
{
return DebugResponse.Failure(request.Command, $"Unknown breakpoint kind '{kindText}'.");
}
var length = 1UL;
if (request.TryGetUInt64("length", out var requestedLength) && requestedLength > 0)
{
length = requestedLength;
}
var breakpoint = _session.Breakpoints.Add(kind, address, length);
return DebugResponse.Success(request.Command, new Dictionary<string, object?>
{
["breakpoint"] = DescribeBreakpoint(breakpoint),
});
}
private DebugResponse RemoveBreakpoint(DebugRequest request)
{
if (!request.TryGetInt32("id", out var id))
{
return DebugResponse.Failure(request.Command, "Expected an 'id'.");
}
return _session.Breakpoints.Remove(id)
? DebugResponse.Success(request.Command)
: DebugResponse.Failure(request.Command, $"No breakpoint with id {id}.");
}
private DebugResponse EnableBreakpoint(DebugRequest request)
{
if (!request.TryGetInt32("id", out var id))
{
return DebugResponse.Failure(request.Command, "Expected an 'id'.");
}
var enabled = !request.TryGetBool("enabled", out var requested) || requested;
return _session.Breakpoints.SetEnabled(id, enabled)
? DebugResponse.Success(request.Command)
: DebugResponse.Failure(request.Command, $"No breakpoint with id {id}.");
}
private DebugResponse Pause(DebugRequest request)
{
_session.RequestPause();
return DebugResponse.Success(request.Command);
}
private static DebugResponse Simple(DebugRequest request, bool succeeded, string failureMessage)
=> succeeded ? DebugResponse.Success(request.Command) : DebugResponse.Failure(request.Command, failureMessage);
private static DebugResponse NotPaused(DebugRequest request)
=> DebugResponse.Failure(request.Command, "Target is not paused.");
internal static IReadOnlyDictionary<string, object?> DescribeStop(DebugStopEvent stop)
{
var data = new Dictionary<string, object?>
{
["reason"] = stop.Reason.ToString(),
["address"] = FormatAddress(stop.Address),
["frameKind"] = stop.FrameKind.ToString(),
["frameLabel"] = stop.FrameLabel,
["registers"] = DescribeRegisters(stop.Registers),
};
if (stop.Breakpoint is { } breakpoint)
{
data["breakpoint"] = DescribeBreakpoint(breakpoint);
}
if (stop.Result is { } result)
{
data["result"] = result.ToString();
}
if (stop.Detail is { } detail)
{
data["detail"] = detail;
}
if (stop.OpcodeBytes is { } opcodeBytes)
{
data["opcodeBytes"] = opcodeBytes;
}
if (stop.StallInfo is { } stall)
{
data["stall"] = new Dictionary<string, object?>
{
["kind"] = stall.Kind.ToString(),
["nid"] = stall.Nid,
["instructionPointer"] = FormatAddress(stall.InstructionPointer),
["dispatchIndex"] = stall.DispatchIndex,
["argument0"] = FormatAddress(stall.Argument0),
["argument1"] = FormatAddress(stall.Argument1),
["resolved"] = stall.IsResolved,
["library"] = stall.LibraryName,
["function"] = stall.FunctionName,
};
}
return data;
}
private static IReadOnlyDictionary<string, object?> DescribeRegisters(DebugRegisterFile registers)
{
var result = new Dictionary<string, object?>(20);
for (var i = 0; i < 16; i++)
{
result[((CpuRegister)i).ToString().ToLowerInvariant()] = FormatAddress(registers[(CpuRegister)i]);
}
result["rip"] = FormatAddress(registers.Rip);
result["rflags"] = FormatAddress(registers.Rflags);
result["fs_base"] = FormatAddress(registers.FsBase);
result["gs_base"] = FormatAddress(registers.GsBase);
return result;
}
private static IReadOnlyDictionary<string, object?> DescribeBreakpoint(Breakpoint breakpoint)
=> new Dictionary<string, object?>
{
["id"] = breakpoint.Id,
["kind"] = breakpoint.Kind.ToString(),
["address"] = FormatAddress(breakpoint.Address),
["length"] = breakpoint.Length,
["enabled"] = breakpoint.Enabled,
};
private static string FormatAddress(ulong value) => $"0x{value:X16}";
private static bool TryParseRegister(string name, out DebugRegisterId id)
{
var normalized = name.Trim().ToLowerInvariant();
switch (normalized)
{
case "rip":
id = DebugRegisterId.Rip;
return true;
case "rflags":
id = DebugRegisterId.Rflags;
return true;
case "fs_base" or "fsbase":
id = DebugRegisterId.FsBase;
return true;
case "gs_base" or "gsbase":
id = DebugRegisterId.GsBase;
return true;
}
return Enum.TryParse(normalized, ignoreCase: true, out id) && Enum.IsDefined(id);
}
private static bool TryParseBreakpointKind(string text, out BreakpointKind kind)
=> Enum.TryParse(text.Trim(), ignoreCase: true, out kind) && Enum.IsDefined(kind);
private const int MaxMemoryChunk = 64 * 1024;
}
@@ -0,0 +1,152 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using System.Globalization;
using System.Text.Json;
namespace SharpEmu.Debugger.Protocol;
/// <summary>
/// A parsed client request: a <see cref="Command"/> verb plus a bag of named
/// arguments backed by the original JSON. Numeric arguments accept either JSON
/// numbers or <c>"0x"</c>-prefixed hex strings so addresses read naturally on
/// the wire.
/// </summary>
public sealed class DebugRequest
{
private readonly JsonElement _root;
private DebugRequest(string command, JsonElement root)
{
Command = command;
_root = root;
}
/// <summary>The lower-cased command verb.</summary>
public string Command { get; }
/// <summary>
/// Parses a single JSON object into a request. Returns false when the text is
/// not a JSON object or is missing a string <c>command</c> field.
/// </summary>
public static bool TryParse(string json, out DebugRequest request, out string error)
{
request = null!;
error = string.Empty;
try
{
using var document = JsonDocument.Parse(json);
var root = document.RootElement.Clone();
if (root.ValueKind != JsonValueKind.Object)
{
error = "Request must be a JSON object.";
return false;
}
if (!root.TryGetProperty("command", out var commandElement) ||
commandElement.ValueKind != JsonValueKind.String)
{
error = "Request is missing a string 'command'.";
return false;
}
var command = commandElement.GetString() ?? string.Empty;
request = new DebugRequest(command.Trim().ToLowerInvariant(), root);
return true;
}
catch (JsonException ex)
{
error = $"Malformed JSON: {ex.Message}";
return false;
}
}
public bool TryGetString(string name, out string value)
{
if (_root.TryGetProperty(name, out var element) && element.ValueKind == JsonValueKind.String)
{
value = element.GetString() ?? string.Empty;
return true;
}
value = string.Empty;
return false;
}
public bool TryGetUInt64(string name, out ulong value)
{
value = 0;
if (!_root.TryGetProperty(name, out var element))
{
return false;
}
switch (element.ValueKind)
{
case JsonValueKind.Number:
return element.TryGetUInt64(out value);
case JsonValueKind.String:
return TryParseNumber(element.GetString(), out value);
default:
return false;
}
}
public bool TryGetInt32(string name, out int value)
{
value = 0;
if (!_root.TryGetProperty(name, out var element))
{
return false;
}
switch (element.ValueKind)
{
case JsonValueKind.Number:
return element.TryGetInt32(out value);
case JsonValueKind.String when TryParseNumber(element.GetString(), out var parsed) && parsed <= int.MaxValue:
value = (int)parsed;
return true;
default:
return false;
}
}
public bool TryGetBool(string name, out bool value)
{
value = false;
if (!_root.TryGetProperty(name, out var element))
{
return false;
}
switch (element.ValueKind)
{
case JsonValueKind.True:
value = true;
return true;
case JsonValueKind.False:
value = false;
return true;
default:
return false;
}
}
private static bool TryParseNumber(string? text, out ulong value)
{
value = 0;
if (string.IsNullOrWhiteSpace(text))
{
return false;
}
text = text.Trim();
if (text.StartsWith("0x", StringComparison.OrdinalIgnoreCase))
{
return ulong.TryParse(text.AsSpan(2), NumberStyles.HexNumber, CultureInfo.InvariantCulture, out value);
}
return ulong.TryParse(text, NumberStyles.Integer, CultureInfo.InvariantCulture, out value);
}
}
@@ -0,0 +1,34 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
namespace SharpEmu.Debugger.Protocol;
/// <summary>
/// The reply to a <see cref="DebugRequest"/>: either success with an optional
/// data payload, or a failure with a human-readable message.
/// </summary>
public sealed class DebugResponse
{
private DebugResponse(bool ok, string? command, IReadOnlyDictionary<string, object?>? data, string? error)
{
Ok = ok;
Command = command;
Data = data;
Error = error;
}
public bool Ok { get; }
/// <summary>Echoes the command the reply answers, when known.</summary>
public string? Command { get; }
public IReadOnlyDictionary<string, object?>? Data { get; }
public string? Error { get; }
public static DebugResponse Success(string command, IReadOnlyDictionary<string, object?>? data = null)
=> new(ok: true, command, data, error: null);
public static DebugResponse Failure(string command, string error)
=> new(ok: false, command, data: null, error);
}
@@ -0,0 +1,33 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
namespace SharpEmu.Debugger.Protocol;
/// <summary>
/// Frames debugger traffic on a connection. A protocol turns bytes into
/// <see cref="DebugRequest"/> objects and serialises <see cref="DebugResponse"/>
/// replies plus asynchronous events (stops, resumes, termination) back to the
/// client. Swapping the implementation (line-delimited JSON today, a GDB remote
/// serial stub later) leaves the session and server untouched.
/// </summary>
public interface IDebugProtocol
{
/// <summary>A short protocol name reported in the handshake.</summary>
string Name { get; }
/// <summary>
/// Reads the next request, or null at end of stream. Parse failures are
/// surfaced as a request with a reserved error command rather than throwing.
/// </summary>
Task<DebugRequest?> ReadRequestAsync(TextReader reader, CancellationToken cancellationToken);
/// <summary>Writes a reply to a request.</summary>
Task WriteResponseAsync(TextWriter writer, DebugResponse response, CancellationToken cancellationToken);
/// <summary>Writes an unsolicited event (for example a stop notification).</summary>
Task WriteEventAsync(
TextWriter writer,
string eventName,
IReadOnlyDictionary<string, object?> data,
CancellationToken cancellationToken);
}
@@ -0,0 +1,112 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using System.Text.Json;
namespace SharpEmu.Debugger.Protocol;
/// <summary>
/// A newline-delimited JSON protocol: one JSON object per line in each
/// direction. Requests carry a <c>command</c>; replies carry <c>ok</c> plus
/// <c>data</c>/<c>error</c>; events carry an <c>event</c> name. It is trivial to
/// drive from a socket, <c>nc</c>, or a small script, which suits bring-up and
/// tooling while a richer protocol is layered on later.
/// </summary>
public sealed class JsonLineDebugProtocol : IDebugProtocol
{
/// <summary>The command assigned to a request that failed to parse.</summary>
public const string ParseErrorCommand = "$parse-error";
private static readonly JsonSerializerOptions SerializerOptions = new()
{
WriteIndented = false,
};
public string Name => "json-lines/1";
public async Task<DebugRequest?> ReadRequestAsync(TextReader reader, CancellationToken cancellationToken)
{
while (true)
{
cancellationToken.ThrowIfCancellationRequested();
var line = await reader.ReadLineAsync(cancellationToken).ConfigureAwait(false);
if (line is null)
{
return null;
}
if (string.IsNullOrWhiteSpace(line))
{
continue;
}
if (DebugRequest.TryParse(line, out var request, out var error))
{
return request;
}
// Surface the parse failure as a synthetic request so the connection
// loop can reply with an error rather than dropping the client.
var envelope = $"{{\"command\":\"{ParseErrorCommand}\",\"message\":{JsonSerializer.Serialize(error)}}}";
if (DebugRequest.TryParse(envelope, out var errorRequest, out _))
{
return errorRequest;
}
}
}
public async Task WriteResponseAsync(TextWriter writer, DebugResponse response, CancellationToken cancellationToken)
{
var payload = new Dictionary<string, object?>
{
["ok"] = response.Ok,
};
if (response.Command is not null)
{
payload["command"] = response.Command;
}
if (response.Data is not null)
{
payload["data"] = response.Data;
}
if (response.Error is not null)
{
payload["error"] = response.Error;
}
await WriteLineAsync(writer, payload, cancellationToken).ConfigureAwait(false);
}
public async Task WriteEventAsync(
TextWriter writer,
string eventName,
IReadOnlyDictionary<string, object?> data,
CancellationToken cancellationToken)
{
var payload = new Dictionary<string, object?>(data.Count + 1)
{
["event"] = eventName,
};
foreach (var (key, value) in data)
{
payload[key] = value;
}
await WriteLineAsync(writer, payload, cancellationToken).ConfigureAwait(false);
}
private static async Task WriteLineAsync(
TextWriter writer,
IReadOnlyDictionary<string, object?> payload,
CancellationToken cancellationToken)
{
cancellationToken.ThrowIfCancellationRequested();
var json = JsonSerializer.Serialize(payload, SerializerOptions);
await writer.WriteLineAsync(json.AsMemory(), cancellationToken).ConfigureAwait(false);
await writer.FlushAsync(cancellationToken).ConfigureAwait(false);
}
}
@@ -0,0 +1,158 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using System.Net.Sockets;
using System.Text;
using SharpEmu.Debugger.Protocol;
using SharpEmu.Debugger.Session;
using SharpEmu.Logging;
namespace SharpEmu.Debugger.Server;
/// <summary>
/// Services a single connected client: reads requests, dispatches them against
/// the shared session, and pushes session lifecycle events. Writes from the
/// request loop and from event callbacks are serialised through one lock so the
/// two never interleave a half-written line.
/// </summary>
internal sealed class DebuggerClientConnection : IAsyncDisposable
{
private static readonly SharpEmuLogger Log = SharpEmuLog.For("SharpEmu.Debugger");
private static readonly UTF8Encoding Utf8NoBom = new(encoderShouldEmitUTF8Identifier: false);
private readonly TcpClient _client;
private readonly IDebuggerSession _session;
private readonly IDebugProtocol _protocol;
private readonly DebugCommandDispatcher _dispatcher;
private readonly SemaphoreSlim _writeLock = new(1, 1);
private TextWriter? _writer;
private CancellationToken _cancellationToken;
public DebuggerClientConnection(TcpClient client, IDebuggerSession session, IDebugProtocol protocol)
{
_client = client;
_session = session;
_protocol = protocol;
_dispatcher = new DebugCommandDispatcher(session);
}
public async Task RunAsync(CancellationToken cancellationToken)
{
_cancellationToken = cancellationToken;
var endpoint = _client.Client.RemoteEndPoint?.ToString() ?? "unknown";
Log.Info($"Debugger client connected: {endpoint}");
using var stream = _client.GetStream();
using var reader = new StreamReader(stream, Utf8NoBom);
await using var writer = new StreamWriter(stream, Utf8NoBom) { AutoFlush = false };
_writer = writer;
_session.Stopped += OnStopped;
_session.Resumed += OnResumed;
_session.Terminated += OnTerminated;
try
{
await SendEventAsync("hello", new Dictionary<string, object?>
{
["protocol"] = _protocol.Name,
["state"] = _session.State.ToString(),
}).ConfigureAwait(false);
while (!cancellationToken.IsCancellationRequested)
{
var request = await _protocol.ReadRequestAsync(reader, cancellationToken).ConfigureAwait(false);
if (request is null)
{
break;
}
var response = _dispatcher.Dispatch(request);
await WriteResponseAsync(response).ConfigureAwait(false);
}
}
catch (OperationCanceledException)
{
// Server shutting down.
}
catch (IOException)
{
// Client dropped the connection.
}
catch (Exception ex)
{
Log.Warn($"Debugger client error ({endpoint}): {ex.Message}");
}
finally
{
_session.Stopped -= OnStopped;
_session.Resumed -= OnResumed;
_session.Terminated -= OnTerminated;
_writer = null;
Log.Info($"Debugger client disconnected: {endpoint}");
}
}
private void OnStopped(object? sender, DebugStopEvent stop)
=> _ = SendEventAsync("stopped", DebugCommandDispatcher.DescribeStop(stop));
private void OnResumed(object? sender, EventArgs e)
=> _ = SendEventAsync("resumed", EmptyData);
private void OnTerminated(object? sender, EventArgs e)
=> _ = SendEventAsync("terminated", EmptyData);
private async Task WriteResponseAsync(DebugResponse response)
{
var writer = _writer;
if (writer is null)
{
return;
}
await _writeLock.WaitAsync(_cancellationToken).ConfigureAwait(false);
try
{
await _protocol.WriteResponseAsync(writer, response, _cancellationToken).ConfigureAwait(false);
}
finally
{
_writeLock.Release();
}
}
private async Task SendEventAsync(string name, IReadOnlyDictionary<string, object?> data)
{
var writer = _writer;
if (writer is null)
{
return;
}
try
{
await _writeLock.WaitAsync(_cancellationToken).ConfigureAwait(false);
try
{
await _protocol.WriteEventAsync(writer, name, data, _cancellationToken).ConfigureAwait(false);
}
finally
{
_writeLock.Release();
}
}
catch (Exception ex) when (ex is IOException or OperationCanceledException or ObjectDisposedException)
{
// The client went away between the event firing and the write.
}
}
public ValueTask DisposeAsync()
{
_writeLock.Dispose();
_client.Dispose();
return ValueTask.CompletedTask;
}
private static readonly IReadOnlyDictionary<string, object?> EmptyData = new Dictionary<string, object?>();
}
@@ -0,0 +1,136 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using System.Collections.Concurrent;
using System.Net;
using System.Net.Sockets;
using SharpEmu.Debugger.Protocol;
using SharpEmu.Debugger.Session;
using SharpEmu.Logging;
namespace SharpEmu.Debugger.Server;
/// <summary>
/// A TCP server that exposes an <see cref="IDebuggerSession"/> to remote
/// clients over a pluggable <see cref="IDebugProtocol"/>. Every connection sees
/// the same session, so multiple clients (for example a UI and a scripted
/// probe) observe a consistent view of the target.
/// </summary>
public sealed class DebuggerServer : IDebuggerServer
{
private static readonly SharpEmuLogger Log = SharpEmuLog.For("SharpEmu.Debugger");
private readonly IDebuggerSession _session;
private readonly DebuggerServerOptions _options;
private readonly Func<IDebugProtocol> _protocolFactory;
private readonly ConcurrentDictionary<DebuggerClientConnection, Task> _connections = new();
private readonly CancellationTokenSource _shutdown = new();
private TcpListener? _listener;
private Task? _acceptLoop;
public DebuggerServer(
IDebuggerSession session,
DebuggerServerOptions? options = null,
Func<IDebugProtocol>? protocolFactory = null)
{
_session = session ?? throw new ArgumentNullException(nameof(session));
_options = options ?? new DebuggerServerOptions();
_protocolFactory = protocolFactory ?? (static () => new JsonLineDebugProtocol());
}
public bool IsListening => _listener is not null;
public IPEndPoint? Endpoint { get; private set; }
public void Start()
{
if (_listener is not null)
{
return;
}
var listener = new TcpListener(_options.BindAddress, _options.Port);
listener.Start(_options.MaxClients);
_listener = listener;
Endpoint = (IPEndPoint?)listener.LocalEndpoint;
Log.Info($"Debug server listening on {Endpoint} (protocol {_protocolFactory().Name})");
_acceptLoop = Task.Run(() => AcceptLoopAsync(listener, _shutdown.Token));
}
private async Task AcceptLoopAsync(TcpListener listener, CancellationToken cancellationToken)
{
while (!cancellationToken.IsCancellationRequested)
{
TcpClient client;
try
{
client = await listener.AcceptTcpClientAsync(cancellationToken).ConfigureAwait(false);
}
catch (OperationCanceledException)
{
break;
}
catch (SocketException) when (cancellationToken.IsCancellationRequested)
{
break;
}
catch (ObjectDisposedException)
{
break;
}
var connection = new DebuggerClientConnection(client, _session, _protocolFactory());
var task = Task.Run(() => ServeAsync(connection, cancellationToken), cancellationToken);
_connections[connection] = task;
}
}
private async Task ServeAsync(DebuggerClientConnection connection, CancellationToken cancellationToken)
{
try
{
await connection.RunAsync(cancellationToken).ConfigureAwait(false);
}
finally
{
_connections.TryRemove(connection, out _);
await connection.DisposeAsync().ConfigureAwait(false);
}
}
public async Task StopAsync()
{
if (_listener is null)
{
return;
}
await _shutdown.CancelAsync().ConfigureAwait(false);
_listener.Stop();
_listener = null;
try
{
if (_acceptLoop is not null)
{
await _acceptLoop.ConfigureAwait(false);
}
await Task.WhenAll(_connections.Values).ConfigureAwait(false);
}
catch (Exception ex) when (ex is OperationCanceledException or SocketException or ObjectDisposedException)
{
// Expected while tearing connections down.
}
_connections.Clear();
Log.Info("Debug server stopped.");
}
public async ValueTask DisposeAsync()
{
await StopAsync().ConfigureAwait(false);
_shutdown.Dispose();
}
}
@@ -0,0 +1,78 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using System.Net;
namespace SharpEmu.Debugger.Server;
/// <summary>Network configuration for a <see cref="DebuggerServer"/>.</summary>
public sealed class DebuggerServerOptions
{
/// <summary>The default TCP port the debug server listens on.</summary>
public const int DefaultPort = 5714;
/// <summary>
/// The address to bind. Defaults to loopback so the debug surface is not
/// exposed off-box; a caller must opt in to a routable address explicitly.
/// </summary>
public IPAddress BindAddress { get; init; } = IPAddress.Loopback;
/// <summary>The TCP port to listen on.</summary>
public int Port { get; init; } = DefaultPort;
/// <summary>
/// The maximum number of simultaneous client connections. Additional
/// connections wait in the accept backlog.
/// </summary>
public int MaxClients { get; init; } = 4;
/// <summary>
/// Parses a <c>host:port</c>, bare <c>port</c>, or bare host into options.
/// Returns false when the text cannot be interpreted.
/// </summary>
public static bool TryParseEndpoint(string? text, out DebuggerServerOptions options, out string error)
{
options = new DebuggerServerOptions();
error = string.Empty;
if (string.IsNullOrWhiteSpace(text))
{
return true;
}
var value = text.Trim();
var host = value;
var port = DefaultPort;
var separator = value.LastIndexOf(':');
if (separator >= 0)
{
var portText = value[(separator + 1)..];
if (portText.Length > 0)
{
if (!int.TryParse(portText, out port) || port is <= 0 or > 65535)
{
error = $"Invalid port '{portText}'.";
return false;
}
}
host = value[..separator];
}
var address = IPAddress.Loopback;
if (!string.IsNullOrWhiteSpace(host) &&
!string.Equals(host, "localhost", StringComparison.OrdinalIgnoreCase) &&
!IPAddress.TryParse(host, out address!))
{
error = $"Invalid bind address '{host}'.";
return false;
}
options = new DebuggerServerOptions
{
BindAddress = address ?? IPAddress.Loopback,
Port = port,
};
return true;
}
}
@@ -0,0 +1,24 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using System.Net;
namespace SharpEmu.Debugger.Server;
/// <summary>
/// A network front-end that exposes a debugger session to remote clients.
/// </summary>
public interface IDebuggerServer : IAsyncDisposable
{
/// <summary>True once the listener is accepting connections.</summary>
bool IsListening { get; }
/// <summary>The endpoint the server is bound to, or null before start.</summary>
IPEndPoint? Endpoint { get; }
/// <summary>Binds and begins accepting client connections.</summary>
void Start();
/// <summary>Stops accepting connections and closes active clients.</summary>
Task StopAsync();
}
@@ -0,0 +1,69 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using SharpEmu.Core.Cpu.Debugging;
using SharpEmu.Debugger.Breakpoints;
using SharpEmu.HLE;
namespace SharpEmu.Debugger.Session;
/// <summary>
/// Describes a stop delivered to debugger clients: why the target stopped,
/// where, and the register snapshot at that point.
/// </summary>
public sealed class DebugStopEvent
{
public DebugStopEvent(
DebugStopReason reason,
DebugRegisterFile registers,
CpuDebugFrameKind frameKind,
string frameLabel,
Breakpoint? breakpoint = null,
OrbisGen2Result? result = null,
string? detail = null,
string? opcodeBytes = null,
CpuStallInfo? stallInfo = null)
{
Reason = reason;
Registers = registers;
FrameKind = frameKind;
FrameLabel = frameLabel ?? string.Empty;
Breakpoint = breakpoint;
Result = result;
Detail = detail;
OpcodeBytes = opcodeBytes;
StallInfo = stallInfo;
}
public DebugStopReason Reason { get; }
/// <summary>The instruction pointer where the target stopped.</summary>
public ulong Address => Registers.Rip;
public DebugRegisterFile Registers { get; }
public CpuDebugFrameKind FrameKind { get; }
public string FrameLabel { get; }
/// <summary>The breakpoint responsible for the stop, when applicable.</summary>
public Breakpoint? Breakpoint { get; }
/// <summary>
/// The frame result for a <see cref="DebugStopReason.Fault"/> stop; null for
/// non-fault stops.
/// </summary>
public OrbisGen2Result? Result { get; }
/// <summary>A human-readable summary of a fault, when applicable.</summary>
public string? Detail { get; }
/// <summary>
/// A hex preview of the bytes at <see cref="Address"/> (the faulting
/// instruction), when the stop is a fault and the bytes were readable.
/// </summary>
public string? OpcodeBytes { get; }
/// <summary>Structured backend evidence for a stall stop.</summary>
public CpuStallInfo? StallInfo { get; }
}
@@ -0,0 +1,32 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
namespace SharpEmu.Debugger.Session;
/// <summary>Why the target stopped and handed control to the debugger.</summary>
public enum DebugStopReason
{
/// <summary>Stopped at the configured entry point before running any frame.</summary>
EntryPoint,
/// <summary>An execution breakpoint was hit.</summary>
Breakpoint,
/// <summary>A data watchpoint was hit.</summary>
Watchpoint,
/// <summary>A single-step (frame step) request completed.</summary>
Step,
/// <summary>A client-requested pause took effect.</summary>
Pause,
/// <summary>The guest raised a fault or trap.</summary>
Fault,
/// <summary>
/// The backend detected an execution stall (for example a mutex spin loop /
/// livelock) with no forward progress.
/// </summary>
Stall,
}
@@ -0,0 +1,23 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
namespace SharpEmu.Debugger.Session;
/// <summary>The execution state of a debugged target as seen by the debugger.</summary>
public enum DebuggerRunState
{
/// <summary>No guest frame has entered the debugger yet.</summary>
Detached,
/// <summary>The guest is executing and cannot be inspected safely.</summary>
Running,
/// <summary>
/// The guest is parked at a frame boundary. Registers and memory can be
/// read and written, and breakpoints can be edited.
/// </summary>
Paused,
/// <summary>The guest has finished; no further frames will run.</summary>
Terminated,
}
@@ -0,0 +1,425 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using SharpEmu.Core.Cpu.Debugging;
using SharpEmu.Debugger.Breakpoints;
using SharpEmu.HLE;
using SharpEmu.Logging;
namespace SharpEmu.Debugger.Session;
/// <summary>
/// The default <see cref="IDebuggerSession"/>. It plugs into the CPU dispatcher
/// as an <see cref="ICpuDebugHook"/>: when a frame boundary warrants a stop it
/// parks the emulation thread inside <see cref="ICpuDebugHook.OnFrameEnter"/>
/// while a debug client inspects and edits state, then releases it on
/// continue/step.
/// </summary>
/// <remarks>
/// Pausing works by blocking the emulation thread on <see cref="_resumeGate"/>
/// from within the hook call. Because that thread is the one that owns the guest
/// context, register and memory accessors are safe to serve from other threads
/// only while it is parked — which is exactly the <see cref="DebuggerRunState.Paused"/>
/// window the accessors gate on.
/// </remarks>
public sealed class DebuggerSession : IDebuggerSession, ICpuDebugHook
{
private static readonly SharpEmuLogger Log = SharpEmuLog.For("SharpEmu.Debugger");
private readonly object _sync = new();
private readonly ManualResetEventSlim _resumeGate = new(initialState: false);
private readonly DebuggerSessionOptions _options;
private ICpuDebugFrame? _currentFrame;
private DebuggerRunState _state = DebuggerRunState.Detached;
private DebugStopEvent? _lastStop;
private bool _seenFirstFrame;
private bool _pausePending;
private bool _stepPending;
public DebuggerSession(DebuggerSessionOptions? options = null)
{
_options = options ?? new DebuggerSessionOptions();
Breakpoints = new BreakpointStore();
}
public BreakpointStore Breakpoints { get; }
public ICpuDebugHook Hook => this;
public event EventHandler<DebugStopEvent>? Stopped;
public event EventHandler? Resumed;
public event EventHandler? Terminated;
public DebuggerRunState State
{
get
{
lock (_sync)
{
return _state;
}
}
}
public DebugStopEvent? LastStop
{
get
{
lock (_sync)
{
return _lastStop;
}
}
}
void ICpuDebugHook.OnFrameEnter(ICpuDebugFrame frame)
{
DebugStopEvent? stop;
lock (_sync)
{
_currentFrame = frame;
var firstFrame = !_seenFirstFrame;
_seenFirstFrame = true;
var reason = ResolveStopReason(frame, firstFrame, out var breakpoint);
if (reason is null)
{
_state = DebuggerRunState.Running;
return;
}
_state = DebuggerRunState.Paused;
_lastStop = new DebugStopEvent(
reason.Value,
DebugRegisterFile.Capture(frame),
frame.Kind,
frame.Label,
breakpoint);
stop = _lastStop;
_resumeGate.Reset();
}
Log.Debug($"Debugger stop: {stop!.Reason} at 0x{stop.Address:X16} ({stop.FrameLabel})");
Stopped?.Invoke(this, stop);
// Park the emulation thread until a client resumes the target. The frame
// stays live and inspectable for the whole wait.
_resumeGate.Wait();
lock (_sync)
{
if (_state != DebuggerRunState.Terminated)
{
_state = DebuggerRunState.Running;
}
}
Resumed?.Invoke(this, EventArgs.Empty);
}
void ICpuDebugHook.OnFrameExit(ICpuDebugFrame frame, OrbisGen2Result result)
{
DebugStopEvent? stop = null;
lock (_sync)
{
if (_options.BreakOnFault &&
result != OrbisGen2Result.ORBIS_GEN2_OK &&
_state != DebuggerRunState.Terminated)
{
// Parking here keeps the post-fault frame inspectable.
_currentFrame = frame;
_state = DebuggerRunState.Paused;
_lastStop = BuildFaultStop(frame, result);
stop = _lastStop;
_resumeGate.Reset();
}
}
if (stop is not null)
{
Log.Debug($"Debugger fault stop: {stop.Result} at 0x{stop.Address:X16} ({stop.FrameLabel})");
Stopped?.Invoke(this, stop);
_resumeGate.Wait();
Resumed?.Invoke(this, EventArgs.Empty);
}
lock (_sync)
{
if (ReferenceEquals(_currentFrame, frame))
{
_currentFrame = null;
}
if (_state != DebuggerRunState.Terminated)
{
_state = DebuggerRunState.Running;
}
}
}
void ICpuDebugHook.OnStall(ICpuDebugFrame frame, CpuStallInfo info)
{
if (!_options.BreakOnStall)
{
return;
}
DebugStopEvent? stop = null;
lock (_sync)
{
if (_state == DebuggerRunState.Terminated)
{
return;
}
_currentFrame = frame;
_state = DebuggerRunState.Paused;
_lastStop = new DebugStopEvent(
DebugStopReason.Stall,
DebugRegisterFile.Capture(frame),
frame.Kind,
frame.Label,
breakpoint: null,
result: null,
detail: info.Detail,
opcodeBytes: ReadOpcodePreview(frame, info.InstructionPointer, 16),
stallInfo: info);
stop = _lastStop;
_resumeGate.Reset();
}
Log.Debug($"Debugger stall stop: {info.Kind} nid={info.Nid} at 0x{info.InstructionPointer:X16}");
Stopped?.Invoke(this, stop);
_resumeGate.Wait();
lock (_sync)
{
if (ReferenceEquals(_currentFrame, frame))
{
_currentFrame = null;
}
if (_state != DebuggerRunState.Terminated)
{
_state = DebuggerRunState.Running;
}
}
Resumed?.Invoke(this, EventArgs.Empty);
}
private static DebugStopEvent BuildFaultStop(ICpuDebugFrame frame, OrbisGen2Result result)
{
var opcodeBytes = ReadOpcodePreview(frame, frame.Rip, 16);
var detail = $"result={result}";
if (opcodeBytes is not null)
{
detail += $", bytes={opcodeBytes}";
}
return new DebugStopEvent(
DebugStopReason.Fault,
DebugRegisterFile.Capture(frame),
frame.Kind,
frame.Label,
breakpoint: null,
result: result,
detail: detail,
opcodeBytes: opcodeBytes);
}
private static string? ReadOpcodePreview(ICpuDebugFrame frame, ulong address, int maxBytes)
{
Span<byte> buffer = stackalloc byte[maxBytes];
var count = 0;
for (; count < maxBytes; count++)
{
if (!frame.Memory.TryRead(address + (ulong)count, buffer.Slice(count, 1)))
{
break;
}
}
return count == 0 ? null : Convert.ToHexString(buffer[..count]);
}
/// <summary>
/// Signals that the whole guest run has finished. Releases any parked
/// emulation thread and moves the session to
/// <see cref="DebuggerRunState.Terminated"/>.
/// </summary>
public void NotifyTerminated()
{
lock (_sync)
{
_state = DebuggerRunState.Terminated;
_currentFrame = null;
}
_resumeGate.Set();
Terminated?.Invoke(this, EventArgs.Empty);
}
private DebugStopReason? ResolveStopReason(ICpuDebugFrame frame, bool firstFrame, out Breakpoint? breakpoint)
{
breakpoint = null;
if (_pausePending)
{
_pausePending = false;
return DebugStopReason.Pause;
}
if (_stepPending)
{
_stepPending = false;
return DebugStopReason.Step;
}
var hit = Breakpoints.FindExecuteHit(frame.EntryPoint);
if (hit is not null)
{
breakpoint = hit;
return DebugStopReason.Breakpoint;
}
if (_options.StopAtEntry && firstFrame)
{
return DebugStopReason.EntryPoint;
}
return null;
}
public bool TryGetRegisters(out DebugRegisterFile registers)
{
lock (_sync)
{
if (!IsPausedWithFrame(out var frame))
{
registers = default;
return false;
}
registers = DebugRegisterFile.Capture(frame);
return true;
}
}
public bool TrySetRegister(DebugRegisterId id, ulong value)
{
lock (_sync)
{
if (!IsPausedWithFrame(out var frame))
{
return false;
}
if (id.IsGeneralPurpose())
{
frame.SetRegister(id.ToCpuRegister(), value);
return true;
}
switch (id)
{
case DebugRegisterId.Rip:
frame.Rip = value;
return true;
case DebugRegisterId.Rflags:
frame.Rflags = value;
return true;
default:
// FS/GS bases are owned by the TLS setup and are read-only here.
return false;
}
}
}
public bool TryReadMemory(ulong address, Span<byte> destination)
{
lock (_sync)
{
return IsPausedWithFrame(out var frame) && frame.Memory.TryRead(address, destination);
}
}
public bool TryWriteMemory(ulong address, ReadOnlySpan<byte> source)
{
lock (_sync)
{
return IsPausedWithFrame(out var frame) && frame.Memory.TryWrite(address, source);
}
}
public bool TryReadXmm(int registerIndex, out ulong low, out ulong high)
{
lock (_sync)
{
if (!IsPausedWithFrame(out var frame) || (uint)registerIndex >= 16)
{
low = 0;
high = 0;
return false;
}
frame.GetXmm(registerIndex, out low, out high);
return true;
}
}
public bool Continue()
{
lock (_sync)
{
if (_state != DebuggerRunState.Paused)
{
return false;
}
_resumeGate.Set();
return true;
}
}
public bool StepFrame()
{
lock (_sync)
{
if (_state != DebuggerRunState.Paused)
{
return false;
}
_stepPending = true;
_resumeGate.Set();
return true;
}
}
public void RequestPause()
{
lock (_sync)
{
if (_state == DebuggerRunState.Running)
{
_pausePending = true;
}
}
}
private bool IsPausedWithFrame(out ICpuDebugFrame frame)
{
// Callers must hold _sync.
if (_state == DebuggerRunState.Paused && _currentFrame is not null)
{
frame = _currentFrame;
return true;
}
frame = null!;
return false;
}
}
@@ -0,0 +1,31 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
namespace SharpEmu.Debugger.Session;
/// <summary>Configuration for a <see cref="DebuggerSession"/>.</summary>
public sealed class DebuggerSessionOptions
{
/// <summary>
/// When true, the session pauses at the first frame it observes so a client
/// can attach breakpoints before the guest runs. Defaults to true, matching
/// the "stop at entry" behaviour most debuggers expose.
/// </summary>
public bool StopAtEntry { get; init; } = true;
/// <summary>
/// When true, the session pauses when a frame ends with a non-OK result (a
/// CPU trap, memory fault, or unimplemented path) so a client can inspect the
/// post-fault register/memory state before the frame is torn down. Defaults
/// to true. The stop reports <see cref="DebugStopReason.Fault"/>.
/// </summary>
public bool BreakOnFault { get; init; } = true;
/// <summary>
/// When true, the session pauses when the backend detects an execution stall
/// (a mutex spin loop / livelock) before the guest is forced out of the loop,
/// so a client can inspect the stalled state. Defaults to true. The stop
/// reports <see cref="DebugStopReason.Stall"/>.
/// </summary>
public bool BreakOnStall { get; init; } = true;
}
@@ -0,0 +1,51 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
namespace SharpEmu.Debugger.Session;
/// <summary>
/// The inspection and control surface a debugger front-end (for example a
/// network server) drives. Register and memory accessors succeed only while the
/// target is <see cref="DebuggerRunState.Paused"/>; they return <c>false</c>
/// otherwise so callers never read torn state from a running guest.
/// </summary>
public interface IDebugTarget
{
/// <summary>The current execution state.</summary>
DebuggerRunState State { get; }
/// <summary>The most recent stop, or null if the target has not stopped yet.</summary>
DebugStopEvent? LastStop { get; }
/// <summary>Reads the integer register file. Fails unless paused.</summary>
bool TryGetRegisters(out DebugRegisterFile registers);
/// <summary>Writes a single register. Fails unless paused.</summary>
bool TrySetRegister(DebugRegisterId id, ulong value);
/// <summary>Reads guest memory into <paramref name="destination"/>. Fails unless paused.</summary>
bool TryReadMemory(ulong address, Span<byte> destination);
/// <summary>Writes guest memory from <paramref name="source"/>. Fails unless paused.</summary>
bool TryWriteMemory(ulong address, ReadOnlySpan<byte> source);
/// <summary>Reads a 128-bit XMM register. Fails unless paused.</summary>
bool TryReadXmm(int registerIndex, out ulong low, out ulong high);
/// <summary>
/// Resumes a paused target. Returns false when the target was not paused.
/// </summary>
bool Continue();
/// <summary>
/// Resumes a paused target and stops again at the next frame boundary.
/// Returns false when the target was not paused.
/// </summary>
bool StepFrame();
/// <summary>
/// Requests that a running target stop at the next frame boundary. Has no
/// effect if the target is already paused or terminated.
/// </summary>
void RequestPause();
}

Some files were not shown because too many files have changed in this diff Show More