Linux and macOS support (#47)

* [macos/linux] Cross-platform host memory, TLS, and ABI layer for POSIX

Introduces the foundation for running SharpEmu on macOS (osx-x64 under
Rosetta 2) and Linux (linux-x64). The CPU backend executes guest x86-64
code natively, so these targets run the whole process as x86-64; this
commit replaces the Windows-only host primitives with platform-dispatched
equivalents so the guest boots and services HLE calls off Windows.

Memory (HostMemory.cs, new): a Win32-semantics facade over
mmap/mprotect/munmap with a shadow region table answering VirtualQuery.
PhysicalVirtualMemory, DirectExecutionBackend, StubManager, and the two
Kernel*CompatExports now go through it instead of kernel32 P/Invokes.
Exact-address requests use MAP_FIXED_NOREPLACE (Linux) / guarded
MAP_FIXED (macOS) so they match Win32 "map there or fail" semantics.

TLS + host helpers (PosixHostStubs.cs, new): pthread-backed TLS and
Win64-ABI-compatible stubs for the kernel32 helpers the backend embeds
into emitted x86-64 code (TlsGetValue, QueryPerformanceCounter,
SwitchToThread, Sleep). A Win64->SysV thunk wraps managed callbacks,
since .NET on POSIX compiles them for the SysV ABI while the emitted
call sites use Win64.

Guest address layout: the 0x7FFx window is Windows-only (dyld shared
cache / Rosetta runtime live there on POSIX), so stack/TLS/stub regions
relocate to 0x6FFx off Windows.

Vectored exception handling is gated off on POSIX for now (guest faults
are not yet recovered) — the signal-based bridge is the next step. Also
adds osx-x64 to the RID list and a Docker-based Linux smoke-test script.

Status: on both macOS (Rosetta) and Linux (amd64), the guest now boots,
runs native x86-64 code, and dispatches HLE imports. macOS stops at a
Rosetta translation-cache issue; Linux runs ~252 imports through C++
static-init before hitting the missing fault handler (SIGSEGV).

* [posix] Bridge the vectored exception handler to sigaction(SIGSEGV/SIGBUS/SIGILL)

Guest faults on macOS/Linux previously terminated the process because the
recovery logic in DirectExecutionBackend.Exceptions.cs was Windows-only.
This adds a POSIX front-end that reuses the existing handler bodies:

- DirectExecutionBackend.PosixSignals.cs installs SA_SIGINFO handlers via
  an [UnmanagedCallersOnly] entry, rebuilds the Win64 EXCEPTION_POINTERS /
  CONTEXT view from the platform mcontext (Darwin __ss thread state via
  the mcontext pointer at ucontext+48, Linux glibc gregs at ucontext+40 --
  offsets verified against the headers on both platforms), runs the same
  chain as the VEH path (TryRecoverUnresolvedSentinel trap-sentinel
  recovery, TryHandleLazyCommittedPage demand paging, VectoredHandler
  diagnostics incl. FS/GS TLS-fault detection), and writes register
  changes back into the mcontext so sigreturn resumes the repaired guest.
  Unrecovered faults chain to the previously installed handler so the
  .NET runtime keeps mapping its own faults to managed exceptions.

- The whole recovery path is warmed up with fabricated inputs before the
  handlers are installed. This is required under Rosetta 2: the signal
  trampoline cannot enter x86 code that has never been executed (and so
  never translated) -- a cold handler is silently never invoked and the
  faulting instruction retries forever (reproduced and verified in an
  isolated .NET test under Rosetta for Linux). It also keeps first-fault
  JIT work out of the signal frame.

- Handlers run without SA_ONSTACK: the runtime's alternate stacks are too
  small for the diagnostic path, while guest (2MB) and host thread stacks
  match where Windows dispatches exceptions anyway.

- The raw reads in the shared fault diagnostics (stack qwords, RBP walk,
  code bytes at RIP) now probe the region table on POSIX before touching
  memory, since a nested SIGSEGV inside the handler would kill the
  process before diagnostics finish. Windows keeps its try/catch reads.

- Escape hatches: SHARPEMU_DISABLE_POSIX_SIGNALS=1 skips installation,
  SHARPEMU_DISABLE_RAW_HANDLER=1 disables sentinel recovery (parity with
  Windows), SHARPEMU_LOG_POSIX_SIGNALS=1 traces every delivery (first 16
  and every 1024th are always traced).

Verified with the test game: Linux (amd64 container) previously died with
SIGSEGV right after import #252; it now recovers/diagnoses signals and the
run proceeds to the real next blocker, an unpatched negative-offset guest
TLS read (fault at TLS base - 0x1708), which gets the full NATIVE
EXCEPTION dump before terminating. macOS is unchanged: the bridge installs
and the game still stops at the known Rosetta translation-cache error at
import 12, which is the next work item.

* [posix] Fix guest memory layout faults: TLS prefix, exact mmap, map search base

Three fixes that take the test game from dying during libc init to running
its full main loop on macOS and Linux:

- Static TLS blocks live below the TCB (FreeBSD amd64 variant II) and
  libc.prx reaches past -0x1700, but only a 4KB prefix was mapped below
  the TLS base. The prefix is now 64KB on POSIX (Windows keeps 4KB); the
  fault was a read at TLS base - 0x1708 during libc init.

- HostMemory exact allocation on macOS used MAP_FIXED, which silently
  maps over untracked host memory. The direct-memory allocator's address
  scan walked into the .NET runtime's JIT heap and replaced live code,
  which under Rosetta 2 surfaced as "no code fragment associated with
  the given arm pc". Exact placement now passes the address as a hint
  and fails on relocation, like MAP_FIXED_NOREPLACE does on Linux.

- sceKernelMapDirectMemory/MapFlexibleMemory searched for free space
  starting at 4GB, which is the Mach-O image base on macOS. The default
  search base is 0x20_0000_0000 on POSIX, and TryAllocateAtOrAbove now
  asks the kernel for a placement instead of page-stepping through host-
  owned address space (Rosetta ignores mmap hints for whole VA windows),
  over-allocating when the caller needs more than page alignment.

Windows behavior is unchanged; all divergences are platform-guarded.

* [macos] Video presenter on the main thread, MoltenVK support, window keyboard input

Gets the test game from a headless loop to a playable window on macOS:

- AppKit traps with SIGILL ("NSUpdateCycleInitialize() is called off the
  main thread") when GLFW runs on a worker thread. The CLI now moves
  emulation onto a worker thread on macOS and parks the real main thread
  in HostMainThread.Pump(); the presenter posts its whole window loop
  there instead of spawning a thread, and a shutdown handler asks the
  render loop to close the window so the pump unwinds on guest exit.

- MoltenVK: enable VK_KHR_portability_enumeration (+ the portability
  instance flag) and VK_KHR_portability_subset when advertised, and gate
  robustBufferAccess2 on the device actually supporting it (Metal does
  not; the old code keyed it off robustImageAccess2 and vkCreateDevice
  failed with ErrorFeatureNotPresent).

- Input: pad exports polled user32 GetAsyncKeyState, so POSIX hosts threw
  DllNotFoundException per scePadReadState call. The presenter now
  attaches the window's keyboard via Silk.NET.Input into HostWindowInput,
  and the pad exports map the existing VK-code layout onto it off
  Windows. Headless hosts (Linux containers) report a disconnected
  keyboard and fall back to neutral pad data silently.

GLFW needs an x86-64 Vulkan loader under Rosetta: place a universal
libMoltenVK.dylib next to SharpEmu named libvulkan.1.dylib (Homebrew's
arm64-only copy cannot load into the x86-64 process) and export
DYLD_LIBRARY_PATH to that directory.

Verified: Dreaming Sarah boots to a MoltenVK-backed 2560x1440 window on
macOS (Apple M4, Rosetta 2), renders the intro, title, and menus, and
keyboard input drives it into gameplay. Linux (amd64 container) runs the
same build headless through millions of imports with no faults. Windows
paths unchanged; arm64 and x64 builds clean.

* [posix] CoreAudio playback, self-contained MoltenVK loading, input/log polish

- Audio: sceAudioOut ports now play through an AudioQueue backend on macOS
  (stereo PCM16 with the same 32KB backpressure pacing as the WinMM path).
  The WinMM port and the new CoreAudio port share an IHostAudioPort
  interface and sample converter; hosts without a backend (Linux
  containers) keep the silent fallback.

- MoltenVK: GLFW resolves Vulkan with dlopen("libvulkan.1.dylib"), which
  cannot see the app-local universal MoltenVK build, so the presenter now
  feeds vkGetInstanceProcAddr straight into glfwInitVulkanLoader (GLFW
  3.4) before creating the window. No DYLD_LIBRARY_PATH needed; the CLI
  also preloads the dylib for Silk.NET and prints setup hints when it is
  missing. scripts/fetch-macos-moltenvk.sh stages the official universal
  dylib next to a build.

- The virtual-range allocator's failure trace now names the address and
  length instead of "AllocateAt invocation threw".

Investigated and documented (not port defects): the savedata transaction
failure is identical on Linux and macOS (HLE argument-register mapping for
sceSaveDataCreateTransactionResource), and the in-game tile speckling has
no platform-specific code in its path - the one macOS-only delta is that
MoltenVK lacks robustBufferAccess2, so out-of-bounds shader reads return
garbage instead of zeros.

Verified on macOS: window, audio backend, and keyboard input all come up
with zero environment configuration; the game runs to gameplay. Linux
headless run unchanged (silent audio, no faults). Windows paths untouched;
arm64 and x64 builds clean.

* [cpu] Preserve guest registers and flags across patched TLS accesses

The TLS patch handler replaces guest `mov reg, fs:[...]` instructions,
which preserve every other register and the flags - but the handler
loaded the TLS index into ecx and called TlsGetValue (Win64: clobbers
rcx/rdx/r8-r11) with `sub/add rsp` trashing the arithmetic flags. Guest
code that keeps live values or comparison results across a TLS access
computed garbage deterministically. The handler now saves rcx, rdx,
r8-r11, and the flags around the call, keeping the same inner stack
alignment. This applies to the load patches and both store-helper stubs,
on every platform.

Also in this change, from the rendering-artifact investigation:

- The present blit picks linear filtering for any fractional scale
  (nearest only for integer upscales): a 3840x2160 guest frame blitted
  into a 2560x1440 swapchain with nearest silently dropped every third
  row/column.
- ClampViewport no longer trims the guest viewport rectangle to the
  render target; trimming changed the guest's scale/offset and skewed
  texel addressing. Vulkan permits viewports beyond the framebuffer
  (the scissor confines rendering), so only spec bounds are enforced.
- Env-gated diagnostics grown during the investigation: guest texture
  dumps (SHARPEMU_TEXTURE_DUMP_DIR), aliased guest-image readback dumps
  (SHARPEMU_TRACE_GUEST_IMAGES=alias), small-render-target write movies
  (SHARPEMU_TRACE_GUEST_WRITES=small), unattended input injection
  (SHARPEMU_AUTO_CROSS=secs,...), viewport nudging
  (SHARPEMU_VIEWPORT_EPSILON), chunked-draw toggle
  (SHARPEMU_DISABLE_CHUNKED_DRAWS), and rect-list/draw vertex traces.

Known remaining issue (root cause narrowed, not yet fixed): the game's
terrain texture pages are corrupted in guest memory before any GPU work
- the mound's solid-fill 32x32 tiles decode to fully transparent texels
and the grass page has deterministic gaps, byte-identical across runs.
Ruled out: memcpy/memmove/memset/realloc HLE semantics, sampler wrap
modes, texel-boundary rounding, chunked draws, viewport handling. Next
step is auditing the Chowdren asset decode path (custom compressed
images) against the emulator's import surface.

* [linux] ALSA playback backend for sceAudioOut

sceAudioOut ports on Linux now play through libasound instead of the
silent fallback. The PCM device opens in blocking mode with ~170ms of
device buffer (the time-equivalent of the 32KB queue the WinMM and
CoreAudio ports keep), so snd_pcm_writei provides the same backpressure
pacing without a managed queue. Underruns and suspend/resume go through
snd_pcm_recover with one retry per submit; anything else drops the
buffer rather than stalling the guest.

The "default" device routes through PulseAudio/PipeWire on desktops
and straight to hardware on bare ALSA; SHARPEMU_ALSA_DEVICE overrides
it (the null device makes the path testable in containers). A missing
libasound or device fails port creation and lands in the existing
silent fallback.

Verified in an amd64 container: the test game opens the port
(backend=alsa, 48kHz stereo float32) and streams sceAudioOutOutput
through the null device for a full run; without a usable device the
port logs a warning and falls back to silent. Playback on real Linux
audio hardware has not been tested.

* [fixes] Address review feedback: commit bounds, CoreAudio shutdown, dump errors

- HostMemory: a MEM_COMMIT that runs past its reservation now fails like
  Win32 instead of committing a prefix and reporting success. All current
  callers already clamp their ranges to the region, so this only guards
  future callers.

- CoreAudioPort: Dispose wakes a submitter waiting on backpressure and
  the wait treats ObjectDisposedException as a timed-out wait, so closing
  a port during playback can no longer throw. A failed AudioQueueStart
  tears the queue down and fails fast instead of leaving an undrainable
  queue that stalls every later submit on its timeout.

- AgcExports: texture dumping catches all write failures (bad path,
  permissions), logging a warning instead of crashing when
  SHARPEMU_TEXTURE_DUMP_DIR points somewhere unusable.

Verified with the Linux container run: game boots and streams audio with
the stricter commit check, and a dump dir under /proc produces warnings
instead of taking the process down.

* [ci] Build linux-x64 and osx-x64 archives

Adds a build-posix matrix job (ubuntu-latest / macos-latest) mirroring
the Windows build: locked restore, Release build, self-contained CLI
publish, and a tar.gz artifact per RID (tar keeps the executable bit).
The macOS archive also stages the universal MoltenVK dylib via
scripts/fetch-macos-moltenvk.sh so the artifact runs without any manual
Vulkan setup. The release job still only ships the Windows archive.

* [cli] Keep POSIX glfw natives outside the single-file bundle

The KeepGlfwOutsideSingleFile target only matched filenames starting
with 'glfw', which covers Windows (glfw3.dll) but not libglfw.3.dylib /
libglfw.so.3. Those got embedded into the single-file bundle, and
Silk.NET's library loader does not probe the bundle extraction
directory, so a published build died with "Couldn't find a suitable
window platform" (and the glfwInitVulkanLoader wiring, which loads the
library from AppContext.BaseDirectory, could not run either). Keeping
the POSIX names loose next to the executable fixes both, the same way
the Windows build already handled it.

Found by running the CI-built osx-x64 archive: video failed while local
loose-file builds worked. With the fix the published single-file build
opens the MoltenVK window, wires the loader, and reaches gameplay.

* [ci] Publish linux-x64 and osx-x64 release archives

The build-posix artifacts now ship as per-RID GitHub releases on main
pushes and manual dispatches, tagged the same way as the win64 ones
(<rid>-<ref>-<sha>). Archives stay tar.gz so the executable bit
survives extraction.

* [cli] Fail early on non-x86-64 host processes

The CPU backend executes guest x86-64 code natively, so the process
must be x86-64 (win-x64/linux-x64 on x64 hardware, osx-x64 under
Rosetta 2 on Apple Silicon). An arm64 process previously failed deep
inside emulation startup, indistinguishable from MoltenVK, signal
handler, or guest memory problems. CLI mode now checks the process
architecture up front and exits with a message naming the supported
execution model (and the Rosetta install command on macOS). The
GUI-only path stays usable on arm64.

* [video] Log the selected Vulkan device name and API version

The presenter never named the GPU it picked, so a 'no video' report
could not be told apart from a real windowing failure without guessing.
It now logs the device name, type, and API version right after
selection. A software rasterizer (llvmpipe/lavapipe/SwiftShader) shows
up here and typically lacks the device features the translated shaders
need, which is the likely cause when a window opens and presents frames
but nothing draws.

* [video] Steer GLFW to XWayland on Wayland sessions

GLFW's native Wayland backend does not reliably map the Vulkan window
with some drivers (NVIDIA in particular): frames present but the window
never becomes visible, so the game runs with audio and no picture. A
report on an RTX 5080 showed exactly this — all device features present,
frames presenting, but the log had 'libdecor-gtk.so failed to init' and
a 1.4x-scaled window, both Wayland tells.

On a Wayland session that also exposes an X server (DISPLAY set), the
presenter now clears WAYLAND_DISPLAY for its own process before GLFW
initializes, so GLFW selects its dependable X11/XWayland backend.
SHARPEMU_ENABLE_WAYLAND=1 opts back into native Wayland. Headless
(no DISPLAY) and non-Linux hosts are unaffected.

* [video] Force GLFW X11 backend via the platform init hint, log the platform

The previous attempt cleared WAYLAND_DISPLAY to steer GLFW off Wayland,
but a reporter still hit the native-Wayland path (the Wayland-only
libdecor error persisted), so that env trick doesn't switch GLFW.

Use GLFW's supported mechanism instead: glfwInitHint(GLFW_PLATFORM,
GLFW_PLATFORM_X11) before GLFW initializes, called into the same libglfw
GLFW itself loads (the pattern InitializeMacVulkanLoader already uses).
Still gated on a Wayland session with an X server present (DISPLAY set)
so we never force X11 where XWayland can't catch it, and still
overridable with SHARPEMU_ENABLE_WAYLAND=1.

Also logs 'GLFW windowing platform in use: <backend>' after init via
glfwGetPlatform, so a 'no window' report shows X11 vs Wayland outright.
Verified on macOS: the readback correctly reports Cocoa and the
presenter is unaffected (the fix is a no-op off Linux).

* [video] Run the GLFW window on the main thread on Linux too

GLFW requires window creation and event processing on the process main
thread on every platform: initialization, window creation, and
glfwPollEvents are main-thread-only, and X11 in particular has a single
event queue that must be serviced there. A window created and polled on
another thread may never map — which is why the game ran (audio, imports,
even Vulkan present) with no visible window on Linux.

macOS already routed the window loop to the main-thread pump (AppKit
needs it); Windows is fine because it has a per-thread event queue. Linux
was the gap: it spawned a background thread for the presenter. Extend the
existing HostMainThread pattern to Linux — emulation runs on a worker,
the main thread pumps the window work the presenter posts.

Refs GLFW intro guide (thread-safety): init, window creation, and event
processing are restricted to the main thread.

Verified: macOS still boots to its window unchanged; the Linux headless
container runs to millions of imports with no deadlock or regression.
On-screen confirmation on a real Linux desktop is still pending, but this
is the documented root cause for a windowless-but-running Linux session.

* [posix] Skip Win32 native guest workers

* [vulkan] Synchronize offscreen targets before present

* [vulkan] Transition fresh textures from undefined layout

* [vulkan] Report swapchain pixels before source readback

* [vulkan] Emit requested guest image diagnostics

* [agc] Diagnose guest texture fallbacks

* [linux] Keep guest GPU mappings in low address space

* [video] Reduce diagnostic stalls and drain complete frames

* [memory] Harden packed GPU address handling

* [readme] Document Linux and macOS release support

* [posix] Integrate the host platform abstraction

* [posix] Restore guest thread address window

* [video] Run the performance HUD on POSIX hosts

The FPS/CPU/work HUD bailed out unless the host was Windows; only the
per-thread hottest-thread scan actually needs Windows APIs. Keep that
scan Windows-only (POSIX reports 'idle') and let the rest of the HUD
run everywhere — the title is already set from the render thread, which
owns the window on macOS and Linux.

* [posix] Implement native guest worker threads

Guest entry stubs must not run above CLR-managed frames on CLR-created
threads (see the NativeWorker preamble); the PR previously fell back to
the inline calli path on POSIX, which reproduced the documented
'attempted to call a UnmanagedCallersOnly method from managed code'
fail-fast (observed after Dreaming Sarah's menu select) and left the
runtime's suspension machinery walking guest frames.

Provide the missing POSIX half of the worker loop:
- PosixHostStubs grows Win64-convention WaitForSingleObject/SetEvent/
  ExitThread stubs backed by dispatch semaphores (macOS) / unnamed POSIX
  semaphores (Linux) plus pthread_exit, with EINTR retry in the wait.
- Worker events are creatable/signalable/waitable from managed code too,
  so NativeGuestExecutor.Run keeps its handshake (AutoResetEvent stays
  on Windows byte-for-byte).
- PosixHostThreading implements CreateNativeThread/WaitForThreadExit/
  CloseThreadHandle over pthreads (liveness probed with
  pthread_kill(0), then joined).
- RunPrologue/RunEpilogue are routed through the existing Win64->SysV
  thunks, so the emitted loop stays identical across platforms.

* [macos] Disable concurrent GC under Rosetta's write-watch hazard

Background GC's write-watch revisit (SoftwareWriteWatch::GetDirty ->
FlushProcessWriteBuffers) calls thread_get_register_pointer_values on
every thread; under Rosetta 2 that Mach call stalls indefinitely on
threads executing translated guest code. The background mark phase then
never finishes and every allocating or Monitor-taking thread wedges
behind it — observed as Dreaming Sarah freezing at the menu/loading
screen with FPS 0 in 5 of 7 runs, dispatcher/watchdog parked in
Monitor.Enter and all BGC threads waiting in t_join.

Non-concurrent GC never takes that path; a 5-minute soak now holds
22-31 fps in-game with zero stalls. Windows and Linux keep concurrent
GC.

* [diag] Periodic guest-thread snapshots with gate-owner tracking

SHARPEMU_PERIODIC_SNAPSHOT_SECONDS=N dumps the stall snapshot every N
seconds even while imports are progressing, for soft stalls where the
game stops advancing but threads keep spinning. The periodic dump never
touches the guest-thread gate (it must keep reporting when the gate is
what's wedged): it reads a lock-free owner record — every gate
acquisition now goes through LockGate(site), which notes site/thread —
and walks the thread table without the lock, tolerating torn reads.
SHARPEMU_PERIODIC_SNAPSHOT_FILE redirects the dump to a side file for
the case where the console itself is wedged (frozen log mirror was one
of the observed failure modes).

* [nuget] Add osx-x64 RID targets to lock files

* [cpu] Back off the guest join poll

TryJoinThread polled the host thread at a fixed 1ms; a game main thread
joining a long-lived worker (Dreaming Sarah parks there for the whole
session) burned ~5% of managed CPU in Join/Sleep syscalls. Ramp the
poll interval to 10ms once the join is clearly long-lived — exit
detection latency for long joins moves from ~1ms to at most 10ms, and
short-lived joins still resolve on the first 1ms polls.

* [nuget] Add linux-x64/win-x64 RID targets to lock files

* [posix] Keep guest stacks clear of the import-stub descent

The import-stub region descends from 0x7000_0000_0000 on the same 16MB
grid as the guest thread windows; moving stacks to 0x6FFF_E000_0000 put
them inside the stub region's 64-module descent range (floor
0x6FFF_C000_0000), silently consuming the top ~32 stack slots on hosts
with many loaded modules. Drop the POSIX stack base to 0x6FFF_A000_0000:
512MB below the stub floor, still 2.5GB above the TLS window. Windows
keeps 0x7FFF_E000_0000 (its bands are ~15TB apart).

* [pad] Read window gamepads on POSIX hosts

XInput and the DualSense hid reader are Windows-only, which left
macOS/Linux with keyboard input only. The presenter's Silk/GLFW input
context already enumerates gamepads on both platforms, so track their
state in HostWindowInput (event-driven on the window thread, snapshot
guarded like the key set) translated to ORBIS conventions: GLFW's Xbox
layout maps A/B/X/Y to Cross/Circle/Square/Triangle, sticks bias from
-1..1 to 0..255 with Y growing down, and triggers rescale from GLFW's
-1..1 resting-at--1 range with digital L2/R2 bits past 25%.

The merge into ReadHostInputState is gated to non-Windows so a physical
pad is never sampled twice through both a native reader and GLFW.
Hotplug is handled via ConnectionChanged; with no pad connected the
path is inert.

Untested against a physical controller (none attached to the dev host);
axis conventions follow the GLFW gamepad-mapping contract.

* [nuget] Refresh lock files after cross-RID restores

* [posix] Adopt the host audio/input seams from main

Main's #192 abstracted audio output and pad/keyboard input behind
IHostAudioOutput/IHostInput; re-express the POSIX backends behind them:

- CoreAudioPort/AlsaAudioPort move to Host/Posix as
  PosixCoreAudioStream/PosixAlsaAudioStream implementing
  IHostAudioStream. The seam converts to stereo PCM16 before Submit, so
  the ports' own conversion (and IHostAudioPort/AudioSampleConverter)
  is gone; queueing and backpressure are unchanged.
- PosixHostAudio selects CoreAudio (macOS) / ALSA (Linux) as the
  platform's IHostAudioOutput.
- PosixHostInput implements IHostInput over an
  IPosixWindowInputSource that HostWindowInput registers when the
  presenter attaches the window's GLFW input context: keyboard with
  virtual-key translation, the window gamepad snapshot (now in seam
  HostGamepadState/HostGamepadButtons terms), and keyboard-connected as
  the focus signal. Rumble/lightbar no-op (GLFW has no such API).
- PadExports drops its direct HostWindowInput gamepad merge — pads now
  flow through IHostInput.GetGamepadStates like every platform.
- PosixHostThreading.RequestTimerResolution is a documented no-op.

All three RIDs build; SharpEmu.Libs.Tests pass (26/26).

* [nuget] Regenerate GUI lock file for RID-less locked restore

Local cross-RID builds stamped a win-x64 runtimes section into
SharpEmu.GUI's lock file; the project declares no RuntimeIdentifiers,
so CI's 'dotnet restore --locked-mode' failed with NU1004 on every
platform. Regenerated via a plain solution restore (--force-evaluate),
matching what the workflow validates.
This commit is contained in:
kuba
2026-07-15 14:36:20 +02:00
committed by GitHub
parent 7b86a91dfa
commit fa2616d224
38 changed files with 4289 additions and 176 deletions
@@ -0,0 +1,173 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using System.Runtime.InteropServices;
namespace SharpEmu.HLE.Host.Posix;
/// <summary>
/// ALSA-based playback for Linux. The PCM device is opened in blocking mode
/// with a device buffer sized to match the 32KB queue the other backends
/// keep, so snd_pcm_writei itself provides the backpressure pacing. The
/// "default" device routes through PulseAudio/PipeWire on desktops and to
/// the hardware on bare ALSA setups; SHARPEMU_ALSA_DEVICE overrides it.
/// </summary>
internal sealed unsafe class PosixAlsaAudioStream : IHostAudioStream
{
// 32KB of stereo PCM16 at 48kHz is ~170ms; keep the same device-side
// queue depth the WinMM/CoreAudio ports enforce in managed code.
private const uint DeviceLatencyMicroseconds = 170_000;
private const int StreamPlayback = 0;
private const int FormatS16LittleEndian = 2;
private const int AccessReadWriteInterleaved = 3;
private const int ErrorPipe = -32; // -EPIPE, underrun
private const int ErrorStreamPipe = -86; // -ESTRPIPE, suspended
private readonly object _gate = new();
private nint _pcm;
private bool _disposed;
public PosixAlsaAudioStream(uint sampleRate)
{
if (!OperatingSystem.IsLinux())
{
throw new PlatformNotSupportedException("ALSA audio is only available on Linux.");
}
var device = Environment.GetEnvironmentVariable("SHARPEMU_ALSA_DEVICE");
if (string.IsNullOrWhiteSpace(device))
{
device = "default";
}
var status = snd_pcm_open(out _pcm, device, StreamPlayback, 0);
if (status != 0)
{
throw new InvalidOperationException(
$"snd_pcm_open(\"{device}\") failed: {DescribeError(status)}.");
}
status = snd_pcm_set_params(
_pcm,
FormatS16LittleEndian,
AccessReadWriteInterleaved,
2,
sampleRate,
1,
DeviceLatencyMicroseconds);
if (status != 0)
{
_ = snd_pcm_close(_pcm);
_pcm = 0;
throw new InvalidOperationException(
$"snd_pcm_set_params({sampleRate} Hz) failed: {DescribeError(status)}.");
}
}
public bool Submit(ReadOnlySpan<byte> stereoPcm16)
{
lock (_gate)
{
if (_disposed)
{
return false;
}
return WritePcm(stereoPcm16, (uint)(stereoPcm16.Length / 4));
}
}
public void Dispose()
{
lock (_gate)
{
if (_disposed)
{
return;
}
_disposed = true;
if (_pcm != 0)
{
_ = snd_pcm_drop(_pcm);
_ = snd_pcm_close(_pcm);
_pcm = 0;
}
}
}
private bool WritePcm(ReadOnlySpan<byte> pcm, uint frames)
{
var recovered = false;
fixed (byte* data = pcm)
{
var offset = 0L;
while (offset < frames)
{
var written = snd_pcm_writei(
_pcm,
data + (offset * 4),
(nuint)(frames - offset));
if (written >= 0)
{
offset += written;
continue;
}
// One recovery attempt per submit covers underruns (-EPIPE)
// and suspend/resume (-ESTRPIPE); anything else, or a second
// failure, drops the buffer rather than stalling the guest.
if (recovered ||
(written != ErrorPipe && written != ErrorStreamPipe) ||
snd_pcm_recover(_pcm, (int)written, 1) != 0)
{
return false;
}
recovered = true;
}
}
return true;
}
private static string DescribeError(long status)
{
var message = Marshal.PtrToStringUTF8(snd_strerror((int)status));
return $"{message ?? "unknown error"} ({status})";
}
private const string Alsa = "libasound.so.2";
[DllImport(Alsa)]
private static extern int snd_pcm_open(
out nint pcm,
[MarshalAs(UnmanagedType.LPUTF8Str)] string name,
int stream,
int mode);
[DllImport(Alsa)]
private static extern int snd_pcm_set_params(
nint pcm,
int format,
int access,
uint channels,
uint rate,
int softResample,
uint latencyUs);
[DllImport(Alsa)]
private static extern long snd_pcm_writei(nint pcm, byte* buffer, nuint frames);
[DllImport(Alsa)]
private static extern int snd_pcm_recover(nint pcm, int error, int silent);
[DllImport(Alsa)]
private static extern int snd_pcm_drop(nint pcm);
[DllImport(Alsa)]
private static extern int snd_pcm_close(nint pcm);
[DllImport(Alsa)]
private static extern nint snd_strerror(int error);
}
@@ -0,0 +1,261 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using System.Runtime.InteropServices;
namespace SharpEmu.HLE.Host.Posix;
/// <summary>
/// AudioQueue-based playback for macOS. Buffers are enqueued as stereo PCM16
/// and returned by the queue's internal thread through the output callback;
/// Submit applies the same 32KB backpressure the WinMM backend uses so guest
/// pacing works identically.
/// </summary>
internal sealed unsafe class PosixCoreAudioStream : IHostAudioStream
{
private const int MaximumQueuedPcmBytes = 32 * 1024;
private const uint FormatLinearPcm = 0x6C70636D; // 'lpcm'
private const uint FlagIsSignedInteger = 0x4;
private const uint FlagIsPacked = 0x8;
private readonly object _gate = new();
private readonly AutoResetEvent _completion = new(false);
private readonly Queue<nint> _freeBuffers = new();
private GCHandle _selfHandle;
private nint _queue;
private int _queuedPcmBytes;
private bool _started;
private bool _disposed;
public PosixCoreAudioStream(uint sampleRate)
{
if (!OperatingSystem.IsMacOS())
{
throw new PlatformNotSupportedException("CoreAudio is only available on macOS.");
}
var format = new AudioStreamBasicDescription
{
SampleRate = sampleRate,
FormatId = FormatLinearPcm,
FormatFlags = FlagIsSignedInteger | FlagIsPacked,
BytesPerPacket = 4,
FramesPerPacket = 1,
BytesPerFrame = 4,
ChannelsPerFrame = 2,
BitsPerChannel = 16,
};
_selfHandle = GCHandle.Alloc(this);
var status = AudioQueueNewOutput(
&format,
&OutputCallback,
GCHandle.ToIntPtr(_selfHandle),
0,
0,
0,
out _queue);
if (status != 0)
{
_selfHandle.Free();
throw new InvalidOperationException($"AudioQueueNewOutput failed with OSStatus {status}.");
}
}
public bool Submit(ReadOnlySpan<byte> stereoPcm16)
{
lock (_gate)
{
if (_disposed || _queue == 0)
{
return false;
}
var outputLength = stereoPcm16.Length;
while (_queuedPcmBytes != 0 &&
_queuedPcmBytes + outputLength > MaximumQueuedPcmBytes)
{
Monitor.Exit(_gate);
try
{
// Dispose can free the event while this thread waits
// outside the gate; treat that like a timed-out wait.
if (!_completion.WaitOne(TimeSpan.FromSeconds(1)))
{
return false;
}
}
catch (ObjectDisposedException)
{
return false;
}
finally
{
Monitor.Enter(_gate);
}
if (_disposed)
{
return false;
}
}
if (!TryTakeBuffer(outputLength, out var buffer))
{
return false;
}
var audioData = ((AudioQueueBuffer*)buffer)->AudioData;
stereoPcm16.CopyTo(new Span<byte>(audioData, outputLength));
((AudioQueueBuffer*)buffer)->AudioDataByteSize = (uint)outputLength;
if (AudioQueueEnqueueBuffer(_queue, buffer, 0, 0) != 0)
{
_freeBuffers.Enqueue(buffer);
return false;
}
_queuedPcmBytes += outputLength;
if (!_started)
{
if (AudioQueueStart(_queue, 0) != 0)
{
// A queue that never starts never drains, so later
// submits would block on backpressure until their
// timeout. Tear the queue down and fail fast instead.
_ = AudioQueueDispose(_queue, true);
_queue = 0;
_queuedPcmBytes = 0;
_freeBuffers.Clear();
return false;
}
_started = true;
}
return true;
}
}
public void Dispose()
{
lock (_gate)
{
if (_disposed)
{
return;
}
_disposed = true;
if (_queue != 0)
{
// Synchronous dispose stops the queue, frees its buffers, and
// guarantees no further callbacks reference this instance.
_ = AudioQueueDispose(_queue, true);
_queue = 0;
}
_freeBuffers.Clear();
// Wake any submitter waiting on backpressure before the event
// goes away; a late waiter observes ObjectDisposedException and
// bails out in Submit.
_completion.Set();
_completion.Dispose();
if (_selfHandle.IsAllocated)
{
_selfHandle.Free();
}
}
}
private bool TryTakeBuffer(int length, out nint buffer)
{
while (_freeBuffers.TryDequeue(out buffer))
{
if (((AudioQueueBuffer*)buffer)->AudioDataBytesCapacity >= (uint)length)
{
return true;
}
_ = AudioQueueFreeBuffer(_queue, buffer);
}
return AudioQueueAllocateBuffer(_queue, (uint)length, out buffer) == 0;
}
[UnmanagedCallersOnly]
private static void OutputCallback(nint userData, nint queue, nint buffer)
{
if (GCHandle.FromIntPtr(userData).Target is not PosixCoreAudioStream port)
{
return;
}
lock (port._gate)
{
if (port._disposed)
{
return;
}
port._queuedPcmBytes -= checked((int)((AudioQueueBuffer*)buffer)->AudioDataByteSize);
port._freeBuffers.Enqueue(buffer);
}
port._completion.Set();
}
[StructLayout(LayoutKind.Sequential)]
private struct AudioStreamBasicDescription
{
public double SampleRate;
public uint FormatId;
public uint FormatFlags;
public uint BytesPerPacket;
public uint FramesPerPacket;
public uint BytesPerFrame;
public uint ChannelsPerFrame;
public uint BitsPerChannel;
public uint Reserved;
}
[StructLayout(LayoutKind.Sequential)]
private struct AudioQueueBuffer
{
public uint AudioDataBytesCapacity;
public void* AudioData;
public uint AudioDataByteSize;
public nint UserData;
public uint PacketDescriptionCapacity;
public nint PacketDescriptions;
public uint PacketDescriptionCount;
}
private const string AudioToolbox =
"/System/Library/Frameworks/AudioToolbox.framework/AudioToolbox";
[DllImport(AudioToolbox)]
private static extern int AudioQueueNewOutput(
AudioStreamBasicDescription* format,
delegate* unmanaged<nint, nint, nint, void> callback,
nint userData,
nint callbackRunLoop,
nint runLoopMode,
uint flags,
out nint queue);
[DllImport(AudioToolbox)]
private static extern int AudioQueueAllocateBuffer(nint queue, uint bufferByteSize, out nint buffer);
[DllImport(AudioToolbox)]
private static extern int AudioQueueFreeBuffer(nint queue, nint buffer);
[DllImport(AudioToolbox)]
private static extern int AudioQueueEnqueueBuffer(nint queue, nint buffer, uint packetDescriptionCount, nint packetDescriptions);
[DllImport(AudioToolbox)]
private static extern int AudioQueueStart(nint queue, nint startTime);
[DllImport(AudioToolbox)]
private static extern int AudioQueueDispose(nint queue, [MarshalAs(UnmanagedType.I1)] bool immediate);
}
@@ -0,0 +1,21 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
namespace SharpEmu.HLE.Host.Posix;
/// <summary>
/// POSIX audio output: CoreAudio (AudioQueue) on macOS, ALSA on Linux. Both
/// streams accept the seam's interleaved stereo PCM16 and pace the guest via
/// device-queue backpressure.
/// </summary>
internal sealed class PosixHostAudio : IHostAudioOutput
{
public string BackendName => OperatingSystem.IsMacOS() ? "coreaudio" : "alsa";
public IHostAudioStream OpenStereoPcm16Stream(uint sampleRate)
{
return OperatingSystem.IsMacOS()
? new PosixCoreAudioStream(sampleRate)
: new PosixAlsaAudioStream(sampleRate);
}
}
@@ -0,0 +1,79 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
namespace SharpEmu.HLE.Host.Posix;
/// <summary>
/// Bridges a window-provided input source into the host input seam. POSIX
/// hosts have no user32/XInput/raw-HID readers; keyboard and gamepad state
/// come from the presenter's GLFW window instead, which registers itself via
/// <see cref="SetSource"/> once the window exists. Until then (and with no
/// window at all, e.g. headless runs) every query reports neutral input.
/// Rumble and lightbar are unsupported by the GLFW input layer and no-op.
/// </summary>
public interface IPosixWindowInputSource
{
/// <summary>True while the window's keyboard is delivering events.</summary>
bool HasKeyboardFocus { get; }
/// <summary>Windows virtual-key semantics; the source translates.</summary>
bool IsKeyDown(int virtualKey);
/// <summary>Same contract as <see cref="IHostInput.GetGamepadStates"/>.</summary>
int GetGamepadStates(Span<HostGamepadState> destination);
string? DescribeConnectedGamepad();
}
// Public so the presenter's window layer (SharpEmu.Libs) can register its
// input source; the platform still constructs the singleton itself.
public sealed class PosixHostInput : IHostInput
{
private static volatile IPosixWindowInputSource? _source;
/// <summary>Called by the presenter's window layer when input is ready.</summary>
public static void SetSource(IPosixWindowInputSource source)
{
_source = source;
}
public void EnsureStarted()
{
// Device readers are event-driven off the window thread; nothing to start.
}
public int GetGamepadStates(Span<HostGamepadState> destination)
{
return _source?.GetGamepadStates(destination) ?? 0;
}
public string? DescribeConnectedGamepad() => _source?.DescribeConnectedGamepad();
public void SetRumble(byte largeMotor, byte smallMotor)
{
}
public void SetTriggerRumble(byte? leftTrigger, byte? rightTrigger)
{
}
public void SetLightbar(byte red, byte green, byte blue)
{
}
public void ResetLightbar()
{
}
public bool IsHostWindowFocused()
{
// GLFW only delivers key events to the focused window, so a
// delivering keyboard implies focus.
return _source?.HasKeyboardFocus ?? false;
}
public bool IsKeyDown(int virtualKey)
{
return _source?.IsKeyDown(virtualKey) ?? false;
}
}
@@ -0,0 +1,508 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using System.Runtime.InteropServices;
namespace SharpEmu.HLE.Host.Posix;
/// <summary>
/// POSIX virtual memory backend implemented over mmap/mprotect/munmap with a
/// shadow region table that answers VirtualQuery-style questions and tracks
/// page protections.
/// POSIX anonymous mappings are demand-paged by the kernel, so Win32
/// "reserve-only" regions are mapped as committed memory directly and
/// commit requests become protection changes.
/// </summary>
internal sealed unsafe class PosixHostMemory : IHostMemory
{
private const uint MEM_COMMIT = 0x1000;
private const uint MEM_RESERVE = 0x2000;
private const uint MEM_RELEASE = 0x8000;
private const uint MEM_FREE_STATE = 0x10000;
private const uint MEM_PRIVATE = 0x20000;
private const uint PAGE_NOACCESS = 0x01;
private const uint PAGE_READONLY = 0x02;
private const uint PAGE_READWRITE = 0x04;
private const uint PAGE_EXECUTE = 0x10;
private const uint PAGE_EXECUTE_READ = 0x20;
private const uint PAGE_EXECUTE_READWRITE = 0x40;
private const ulong PageSize = 0x1000;
private struct BasicInfo
{
public ulong BaseAddress;
public ulong AllocationBase;
public uint AllocationProtect;
public ulong RegionSize;
public uint State;
public uint Protect;
public uint Type;
}
public ulong Allocate(ulong desiredAddress, ulong size, HostPageProtection protection)
{
return (ulong)Posix.Alloc(
(void*)desiredAddress,
(nuint)size,
MEM_COMMIT | MEM_RESERVE,
ToNativeProtection(protection));
}
public ulong Reserve(ulong desiredAddress, ulong size, HostPageProtection protection)
{
return (ulong)Posix.Alloc(
(void*)desiredAddress,
(nuint)size,
MEM_RESERVE,
ToNativeProtection(protection));
}
public bool Commit(ulong address, ulong size, HostPageProtection protection)
{
return Posix.Alloc(
(void*)address,
(nuint)size,
MEM_COMMIT,
ToNativeProtection(protection)) != null;
}
public bool Free(ulong address)
{
return Posix.Free((void*)address, 0, MEM_RELEASE);
}
public bool Protect(
ulong address,
ulong size,
HostPageProtection protection,
out uint rawOldProtection)
{
return Posix.Protect(
(void*)address,
(nuint)size,
ToNativeProtection(protection),
out rawOldProtection);
}
public bool ProtectRaw(
ulong address,
ulong size,
uint rawProtection,
out uint rawOldProtection)
{
return Posix.Protect(
(void*)address,
(nuint)size,
rawProtection,
out rawOldProtection);
}
public bool Query(ulong address, out HostRegionInfo info)
{
if (Posix.Query((void*)address, out var nativeInfo) == 0)
{
info = default;
return false;
}
info = new HostRegionInfo(
nativeInfo.BaseAddress,
nativeInfo.AllocationBase,
nativeInfo.RegionSize,
nativeInfo.State switch
{
MEM_COMMIT => HostRegionState.Committed,
MEM_RESERVE => HostRegionState.Reserved,
_ => HostRegionState.Free,
},
nativeInfo.State,
ToHostProtection(nativeInfo.Protect),
nativeInfo.Protect,
nativeInfo.AllocationProtect);
return true;
}
public void FlushInstructionCache(ulong address, ulong size)
{
_ = address;
_ = size;
// The supported POSIX process is x86-64 (including Rosetta 2), whose
// instruction cache is coherent. A future arm64 backend must call the
// platform instruction-cache invalidation API here.
}
private static uint ToNativeProtection(HostPageProtection protection) => protection switch
{
HostPageProtection.NoAccess => PAGE_NOACCESS,
HostPageProtection.ReadOnly => PAGE_READONLY,
HostPageProtection.ReadWrite => PAGE_READWRITE,
HostPageProtection.Execute => PAGE_EXECUTE,
HostPageProtection.ReadExecute => PAGE_EXECUTE_READ,
HostPageProtection.ReadWriteExecute => PAGE_EXECUTE_READWRITE,
HostPageProtection.ExecuteWriteCopy => PAGE_EXECUTE_READWRITE,
_ => throw new ArgumentOutOfRangeException(nameof(protection), protection, null),
};
private static HostPageProtection ToHostProtection(uint protection) => protection switch
{
PAGE_READONLY => HostPageProtection.ReadOnly,
PAGE_READWRITE => HostPageProtection.ReadWrite,
PAGE_EXECUTE => HostPageProtection.Execute,
PAGE_EXECUTE_READ => HostPageProtection.ReadExecute,
PAGE_EXECUTE_READWRITE => HostPageProtection.ReadWriteExecute,
_ => HostPageProtection.NoAccess,
};
private static class Posix
{
private const int PROT_NONE = 0x0;
private const int PROT_READ = 0x1;
private const int PROT_WRITE = 0x2;
private const int PROT_EXEC = 0x4;
private const int MAP_PRIVATE = 0x02;
private const int MAP_FIXED = 0x10;
private static readonly int MAP_ANON = OperatingSystem.IsMacOS() ? 0x1000 : 0x20;
private static readonly int MAP_NORESERVE = OperatingSystem.IsMacOS() ? 0 : 0x4000;
// Linux-only: fail instead of clobbering an existing mapping.
private const int MAP_FIXED_NOREPLACE = 0x100000;
private static readonly nint MAP_FAILED = -1;
private static readonly object Gate = new();
private static readonly SortedList<ulong, Region> Regions = new();
private sealed class Region
{
public ulong Base;
public ulong Size;
public uint DefaultProtect;
public Dictionary<ulong, uint>? PageProtects;
public ulong End => Base + Size;
public uint ProtectAt(ulong pageAddress)
{
if (PageProtects is not null && PageProtects.TryGetValue(pageAddress, out var overriden))
{
return overriden;
}
return DefaultProtect;
}
}
public static void* Alloc(void* address, nuint size, uint allocationType, uint protect)
{
if (size == 0)
{
return null;
}
var alignedSize = AlignUp((ulong)size, PageSize);
lock (Gate)
{
if (allocationType == MEM_COMMIT && address != null &&
TryFindRegionLocked((ulong)address, out var existing))
{
// Note: MEM_RESERVE requests that overlap an existing
// region must fail like Win32 does; only a pure commit
// may target pages inside a tracked mapping.
// Commit inside an existing mapping: the pages are already
// backed (demand paged), so only apply the protection.
var start = AlignDown((ulong)address, PageSize);
var end = AlignUp((ulong)address + alignedSize, PageSize);
if (end <= start || end > existing.End)
{
// Win32 fails a commit that runs past its reservation
// instead of committing a prefix; committing partially
// here would let callers believe the whole range is
// usable.
return null;
}
if (mprotect((nint)start, (nuint)(end - start), ToPosixProtect(protect)) != 0)
{
return null;
}
SetProtectRangeLocked(existing, start, end - start, protect);
return address;
}
if ((allocationType & MEM_RESERVE) == 0)
{
// MEM_COMMIT alone outside any known region is invalid here.
return null;
}
var posixProtect = ToPosixProtect(protect);
var flags = MAP_PRIVATE | MAP_ANON;
if ((allocationType & MEM_COMMIT) == 0)
{
// Reserve-only: keep the requested protection so the region
// is usable without a separate commit step, but tell the
// kernel not to account swap for it where supported.
flags |= MAP_NORESERVE;
}
nint result;
if (address != null)
{
// Win32 maps at exactly the requested address or fails
// without touching existing mappings. Fail up front on
// any overlap we track, then place the mapping: Linux
// gets MAP_FIXED_NOREPLACE (fails cleanly on host
// mappings too). Darwin lacks NOREPLACE and plain
// MAP_FIXED would silently clobber untracked host
// memory (dyld, the runtime's JIT heap, Rosetta), so
// pass the address as a hint instead -- the kernel
// honors it when the range is free and relocates the
// mapping otherwise, which we treat as failure.
if (OverlapsTrackedRegionLocked((ulong)address, alignedSize))
{
Trace($"exact overlap: addr=0x{(ulong)address:X16} size=0x{alignedSize:X}");
return null;
}
var exactFlags = OperatingSystem.IsMacOS() ? flags : flags | MAP_FIXED_NOREPLACE;
result = mmap((nint)address, (nuint)alignedSize, posixProtect, exactFlags, -1, 0);
if (result == MAP_FAILED || (ulong)result != (ulong)address)
{
Trace($"exact mmap failed: addr=0x{(ulong)address:X16} got=0x{(ulong)result:X16} size=0x{alignedSize:X} errno={Marshal.GetLastPInvokeError()}");
if (result != MAP_FAILED)
{
munmap(result, (nuint)alignedSize);
}
return null;
}
}
else
{
result = mmap(0, (nuint)alignedSize, posixProtect, flags, -1, 0);
if (result == MAP_FAILED)
{
Trace($"mmap failed: size=0x{alignedSize:X} errno={Marshal.GetLastPInvokeError()}");
return null;
}
}
Regions[(ulong)result] = new Region
{
Base = (ulong)result,
Size = alignedSize,
DefaultProtect = protect
};
return (void*)result;
}
}
public static bool Free(void* address, nuint size, uint freeType)
{
_ = size;
_ = freeType;
lock (Gate)
{
if (!Regions.TryGetValue((ulong)address, out var region))
{
return false;
}
Regions.Remove((ulong)address);
return munmap((nint)address, (nuint)region.Size) == 0;
}
}
public static bool Protect(void* address, nuint size, uint newProtect, out uint oldProtect)
{
oldProtect = PAGE_NOACCESS;
if (size == 0)
{
return false;
}
var start = AlignDown((ulong)address, PageSize);
var end = AlignUp((ulong)address + size, PageSize);
lock (Gate)
{
if (!TryFindRegionLocked(start, out var region) || end > region.End)
{
return false;
}
oldProtect = region.ProtectAt(start);
if (mprotect((nint)start, (nuint)(end - start), ToPosixProtect(newProtect)) != 0)
{
return false;
}
SetProtectRangeLocked(region, start, end - start, newProtect);
return true;
}
}
public static nuint Query(void* address, out BasicInfo info)
{
info = default;
var pageAddress = AlignDown((ulong)address, PageSize);
lock (Gate)
{
if (TryFindRegionLocked(pageAddress, out var region))
{
// Win32 VirtualQuery reports a run of pages sharing the
// same protection, so stop the run where it changes.
var protect = region.ProtectAt(pageAddress);
var runEnd = pageAddress + PageSize;
while (runEnd < region.End && region.ProtectAt(runEnd) == protect)
{
runEnd += PageSize;
}
info.BaseAddress = pageAddress;
info.AllocationBase = region.Base;
info.AllocationProtect = region.DefaultProtect;
info.RegionSize = runEnd - pageAddress;
info.State = MEM_COMMIT;
info.Protect = protect;
info.Type = MEM_PRIVATE;
return (nuint)sizeof(BasicInfo);
}
// Untracked host memory (runtime heaps, stacks, libraries) is
// reported as a free block reaching to the next tracked region
// so scanning callers keep advancing.
var nextBase = ulong.MaxValue;
foreach (var regionBase in Regions.Keys)
{
if (regionBase > pageAddress)
{
nextBase = regionBase;
break;
}
}
info.BaseAddress = pageAddress;
info.AllocationBase = 0;
info.AllocationProtect = PAGE_NOACCESS;
info.RegionSize = (nextBase == ulong.MaxValue ? pageAddress + PageSize : nextBase) - pageAddress;
info.State = MEM_FREE_STATE;
info.Protect = PAGE_NOACCESS;
info.Type = 0;
return (nuint)sizeof(BasicInfo);
}
}
private static bool OverlapsTrackedRegionLocked(ulong start, ulong size)
{
var end = start + size;
foreach (var region in Regions.Values)
{
if (region.Base < end && start < region.End)
{
return true;
}
}
return false;
}
private static bool TryFindRegionLocked(ulong address, out Region region)
{
region = null!;
var keys = Regions.Keys;
var low = 0;
var high = keys.Count - 1;
Region? candidate = null;
while (low <= high)
{
var middle = low + ((high - low) >> 1);
var entry = Regions.Values[middle];
if (entry.Base <= address)
{
candidate = entry;
low = middle + 1;
}
else
{
high = middle - 1;
}
}
if (candidate is null || address >= candidate.End)
{
return false;
}
region = candidate;
return true;
}
private static void SetProtectRangeLocked(Region region, ulong start, ulong size, uint protect)
{
if (start == region.Base && size >= region.Size)
{
region.DefaultProtect = protect;
region.PageProtects = null;
return;
}
region.PageProtects ??= new Dictionary<ulong, uint>();
var end = start + size;
for (var pageAddress = start; pageAddress < end; pageAddress += PageSize)
{
if (protect == region.DefaultProtect)
{
region.PageProtects.Remove(pageAddress);
}
else
{
region.PageProtects[pageAddress] = protect;
}
}
}
private static int ToPosixProtect(uint win32Protect)
{
return win32Protect switch
{
PAGE_NOACCESS => PROT_NONE,
PAGE_READONLY => PROT_READ,
PAGE_READWRITE => PROT_READ | PROT_WRITE,
PAGE_EXECUTE => PROT_READ | PROT_EXEC,
PAGE_EXECUTE_READ => PROT_READ | PROT_EXEC,
PAGE_EXECUTE_READWRITE => PROT_READ | PROT_WRITE | PROT_EXEC,
_ => PROT_READ | PROT_WRITE
};
}
private static void Trace(string message)
{
if (string.Equals(Environment.GetEnvironmentVariable("SHARPEMU_LOG_VMEM"), "1", StringComparison.Ordinal))
{
Console.Error.WriteLine($"[HOSTMEM] {message}");
}
}
private static ulong AlignDown(ulong value, ulong alignment) => value & ~(alignment - 1);
private static ulong AlignUp(ulong value, ulong alignment) => checked((value + alignment - 1) & ~(alignment - 1));
[DllImport("libc", SetLastError = true)]
private static extern nint mmap(nint addr, nuint length, int prot, int flags, int fd, long offset);
[DllImport("libc", SetLastError = true)]
private static extern int munmap(nint addr, nuint length);
[DllImport("libc", SetLastError = true)]
private static extern int mprotect(nint addr, nuint length, int prot);
}
}
@@ -0,0 +1,17 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
namespace SharpEmu.HLE.Host.Posix;
internal sealed class PosixHostPlatform : IHostPlatform
{
public IHostMemory Memory { get; } = new PosixHostMemory();
public IHostThreading Threading { get; } = new PosixHostThreading();
public IHostSymbolResolver Symbols { get; } = new PosixHostSymbolResolver();
public IHostAudioOutput Audio { get; } = new PosixHostAudio();
public IHostInput Input { get; } = new PosixHostInput();
}
@@ -0,0 +1,657 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
using System.Runtime.InteropServices;
namespace SharpEmu.HLE.Host.Posix;
/// <summary>
/// POSIX replacements for the kernel32 helpers the native backend embeds in
/// emitted x86-64 code. Every stub exposed here follows the Win64 calling
/// convention the emitted call sites were written for (first argument in
/// ECX, result in RAX, Win64 non-volatile registers preserved), so the
/// emission code stays identical across platforms.
/// </summary>
internal static unsafe class PosixHostStubs
{
private static readonly object Gate = new();
private static bool _initialized;
private static nint _tlsGetValueStub;
private static nint _queryPerformanceCounterStub;
private static nint _switchToThreadStub;
private static nint _sleepStub;
private static nint _waitForSingleObjectStub;
private static nint _setEventStub;
private static nint _exitThreadStub;
public static nint TlsGetValueStubAddress
{
get { EnsureInitialized(); return _tlsGetValueStub; }
}
public static nint QueryPerformanceCounterStubAddress
{
get { EnsureInitialized(); return _queryPerformanceCounterStub; }
}
public static nint SwitchToThreadStubAddress
{
get { EnsureInitialized(); return _switchToThreadStub; }
}
public static nint SleepStubAddress
{
get { EnsureInitialized(); return _sleepStub; }
}
/// <summary>
/// Win64-convention replacements for the kernel32 event/thread helpers the
/// native guest worker loop embeds. The "handle" they take is a worker
/// event created by <see cref="CreateWorkerEvent"/>: a dispatch semaphore
/// on macOS, an unnamed POSIX semaphore on Linux. The wait stub always
/// waits forever (the worker loop passes INFINITE) and retries EINTR.
/// </summary>
public static nint WaitForSingleObjectStubAddress
{
get { EnsureInitialized(); return _waitForSingleObjectStub; }
}
public static nint SetEventStubAddress
{
get { EnsureInitialized(); return _setEventStub; }
}
public static nint ExitThreadStubAddress
{
get { EnsureInitialized(); return _exitThreadStub; }
}
/// <summary>
/// Creates a binary-semaphore worker event signalable/waitable both from
/// managed code and from emitted native code (via the stub addresses
/// above). Returns 0 on failure.
/// </summary>
public static nint CreateWorkerEvent()
{
if (OperatingSystem.IsMacOS())
{
return dispatch_semaphore_create(0);
}
var semaphore = Marshal.AllocHGlobal(64);
if (sem_init(semaphore, 0, 0) != 0)
{
Marshal.FreeHGlobal(semaphore);
return 0;
}
return semaphore;
}
public static bool SignalWorkerEvent(nint handle)
{
if (OperatingSystem.IsMacOS())
{
_ = dispatch_semaphore_signal(handle);
return true;
}
return sem_post(handle) == 0;
}
/// <summary>Waits for a worker event; a negative timeout waits forever.</summary>
public static bool WaitWorkerEvent(nint handle, int timeoutMilliseconds)
{
if (OperatingSystem.IsMacOS())
{
if (timeoutMilliseconds < 0)
{
return dispatch_semaphore_wait(handle, ulong.MaxValue) == 0;
}
var deadline = dispatch_time(0, timeoutMilliseconds * 1_000_000L);
return dispatch_semaphore_wait(handle, deadline) == 0;
}
if (timeoutMilliseconds < 0)
{
while (sem_wait(handle) != 0)
{
// EINTR: retry.
}
return true;
}
var deadlineTicks = Environment.TickCount64 + timeoutMilliseconds;
while (sem_trywait(handle) != 0)
{
if (Environment.TickCount64 >= deadlineTicks)
{
return false;
}
System.Threading.Thread.Sleep(1);
}
return true;
}
public static void DestroyWorkerEvent(nint handle)
{
if (handle == 0)
{
return;
}
if (OperatingSystem.IsMacOS())
{
dispatch_release(handle);
return;
}
_ = sem_destroy(handle);
Marshal.FreeHGlobal(handle);
}
/// <summary>
/// Starts a raw pthread at a native entry point (pthread entries take their
/// argument in RDI; the worker loop stub ignores it). Returns an opaque
/// handle for <see cref="WaitForWorkerThreadExit"/>/<see cref="CloseWorkerThreadHandle"/>,
/// or 0 on failure.
/// </summary>
public static nint CreateWorkerThread(nint entry, nint parameter, nuint stackReserveBytes, out uint threadId)
{
threadId = 0;
byte* attr = stackalloc byte[512];
if (pthread_attr_init(attr) != 0)
{
return 0;
}
try
{
if (stackReserveBytes != 0)
{
_ = pthread_attr_setstacksize(attr, nuint.Max(stackReserveBytes, 512 * 1024));
}
nint thread;
if (pthread_create(&thread, attr, entry, parameter) != 0)
{
return 0;
}
if (OperatingSystem.IsMacOS())
{
ulong numericId;
if (pthread_threadid_np(thread, &numericId) == 0)
{
threadId = unchecked((uint)numericId);
}
}
else
{
threadId = unchecked((uint)thread);
}
var holder = (nint*)Marshal.AllocHGlobal(sizeof(nint) * 2);
holder[0] = thread;
holder[1] = 0; // joined flag
return (nint)holder;
}
finally
{
_ = pthread_attr_destroy(attr);
}
}
/// <summary>
/// Waits for a worker thread to exit. Liveness is probed with
/// pthread_kill(thread, 0) (ESRCH once the thread has terminated) because
/// neither platform offers a portable timed join; the exited thread is then
/// joined so its resources are reclaimed.
/// </summary>
public static bool WaitForWorkerThreadExit(nint threadHandle, uint timeoutMilliseconds)
{
var holder = (nint*)threadHandle;
if (holder == null)
{
return false;
}
if (holder[1] != 0)
{
return true;
}
var thread = holder[0];
var deadline = Environment.TickCount64 + timeoutMilliseconds;
while (pthread_kill(thread, 0) == 0)
{
if (Environment.TickCount64 >= deadline)
{
return false;
}
System.Threading.Thread.Sleep(1);
}
_ = pthread_join(thread, null);
holder[1] = 1;
return true;
}
public static void CloseWorkerThreadHandle(nint threadHandle)
{
var holder = (nint*)threadHandle;
if (holder == null)
{
return;
}
if (holder[1] == 0)
{
// Never observed exiting: detach so the thread does not leak a
// zombie join target when it eventually terminates.
_ = pthread_detach(holder[0]);
}
Marshal.FreeHGlobal(threadHandle);
}
/// <summary>Allocates a pthread TLS key, mirroring kernel32!TlsAlloc.</summary>
public static uint TlsAlloc()
{
if (OperatingSystem.IsMacOS())
{
nuint key;
return pthread_key_create_mac(&key, 0) == 0 ? (uint)key : uint.MaxValue;
}
uint key32;
return pthread_key_create_linux(&key32, 0) == 0 ? key32 : uint.MaxValue;
}
public static bool TlsFree(uint key)
{
return OperatingSystem.IsMacOS()
? pthread_key_delete_mac((nuint)key) == 0
: pthread_key_delete_linux(key) == 0;
}
public static bool TlsSetValue(uint key, nint value)
{
return OperatingSystem.IsMacOS()
? pthread_setspecific_mac((nuint)key, value) == 0
: pthread_setspecific_linux(key, value) == 0;
}
public static nint TlsGetValue(uint key)
{
return OperatingSystem.IsMacOS()
? pthread_getspecific_mac((nuint)key)
: pthread_getspecific_linux(key);
}
/// <summary>Stable numeric id of the calling thread (kernel32!GetCurrentThreadId).</summary>
public static uint GetCurrentThreadId()
{
if (OperatingSystem.IsMacOS())
{
ulong tid;
return pthread_threadid_np(0, &tid) == 0 ? unchecked((uint)tid) : 0u;
}
return unchecked((uint)gettid());
}
/// <summary>
/// Wraps a managed callback (compiled for the SysV ABI on POSIX .NET) in a
/// thunk that accepts up to four integer arguments in the Win64 ABI the
/// emitted x86-64 call sites use. Win64 passes args in rcx/rdx/r8/r9 and
/// treats rdi/rsi as non-volatile; SysV expects rdi/rsi/rdx/rcx and
/// clobbers them, so the thunk saves rdi/rsi, shuffles the registers, keeps
/// the stack 16-byte aligned for the call, and forwards the rax result.
/// </summary>
public static nint CreateWin64ToSysVThunk(nint sysvTarget)
{
var memory = HostPlatform.Current.Memory;
var page = (byte*)memory.Allocate(
0,
4096,
HostPageProtection.ReadWriteExecute);
if (page == null)
{
throw new OutOfMemoryException("Failed to allocate Win64->SysV thunk page");
}
var offset = 0;
Emit(page, ref offset, 0x57); // push rdi
Emit(page, ref offset, 0x56); // push rsi
Emit(page, ref offset, 0x48, 0x89, 0xCF); // mov rdi, rcx
Emit(page, ref offset, 0x48, 0x89, 0xD6); // mov rsi, rdx
Emit(page, ref offset, 0x4C, 0x89, 0xC2); // mov rdx, r8
Emit(page, ref offset, 0x4C, 0x89, 0xC9); // mov rcx, r9
Emit(page, ref offset, 0x48, 0x83, 0xEC, 0x08); // sub rsp, 8 (realign to 16)
EmitMovRaxImm64(page, ref offset, sysvTarget); // mov rax, target
Emit(page, ref offset, 0xFF, 0xD0); // call rax
Emit(page, ref offset, 0x48, 0x83, 0xC4, 0x08); // add rsp, 8
Emit(page, ref offset, 0x5E); // pop rsi
Emit(page, ref offset, 0x5F); // pop rdi
Emit(page, ref offset, 0xC3); // ret
if (!memory.Protect((ulong)page, 4096, HostPageProtection.ReadExecute, out _))
{
throw new InvalidOperationException("Failed to protect Win64->SysV thunk page");
}
memory.FlushInstructionCache((ulong)page, (ulong)offset);
return (nint)page;
}
private static void EnsureInitialized()
{
if (_initialized)
{
return;
}
lock (Gate)
{
if (_initialized)
{
return;
}
BuildStubs();
_initialized = true;
}
}
private static void BuildStubs()
{
var memory = HostPlatform.Current.Memory;
var page = (byte*)memory.Allocate(
0,
4096,
HostPageProtection.ReadWriteExecute);
if (page == null)
{
throw new OutOfMemoryException("Failed to allocate POSIX host helper stub page");
}
var offset = 0;
_tlsGetValueStub = EmitTlsGetValue(page, ref offset);
_queryPerformanceCounterStub = EmitQueryPerformanceCounter(page, ref offset);
_switchToThreadStub = EmitSwitchToThread(page, ref offset);
_sleepStub = EmitSleep(page, ref offset);
_waitForSingleObjectStub = EmitWaitForSingleObject(page, ref offset);
_setEventStub = EmitSetEvent(page, ref offset);
_exitThreadStub = EmitExitThread(page, ref offset);
if (!memory.Protect((ulong)page, 4096, HostPageProtection.ReadExecute, out _))
{
throw new InvalidOperationException("Failed to protect POSIX host helper stub page");
}
memory.FlushInstructionCache((ulong)page, (ulong)offset);
}
private static nint EmitTlsGetValue(byte* page, ref int offset)
{
var start = (nint)(page + offset);
if (OperatingSystem.IsMacOS())
{
// On macOS x86-64 pthread keys index the gs-based thread specific
// data array directly, so TlsGetValue(index in ecx) collapses to a
// single load that clobbers nothing but RAX.
Emit(page, ref offset, 0x89, 0xC8); // mov eax, ecx
Emit(page, ref offset, 0x65, 0x48, 0x8B, 0x04, 0xC5, 0, 0, 0, 0); // mov rax, gs:[rax*8]
Emit(page, ref offset, 0xC3); // ret
return start;
}
// Linux: call pthread_getspecific, preserving the registers that are
// volatile in SysV but non-volatile in Win64 (rsi, rdi).
var pthreadGetSpecific = ResolveLibcExport("pthread_getspecific");
Emit(page, ref offset, 0x56); // push rsi
Emit(page, ref offset, 0x57); // push rdi
Emit(page, ref offset, 0x48, 0x83, 0xEC, 0x08); // sub rsp, 8
Emit(page, ref offset, 0x89, 0xCF); // mov edi, ecx
EmitMovRaxImm64(page, ref offset, pthreadGetSpecific); // mov rax, imm64
Emit(page, ref offset, 0xFF, 0xD0); // call rax
Emit(page, ref offset, 0x48, 0x83, 0xC4, 0x08); // add rsp, 8
Emit(page, ref offset, 0x5F); // pop rdi
Emit(page, ref offset, 0x5E); // pop rsi
Emit(page, ref offset, 0xC3); // ret
return start;
}
private static nint EmitQueryPerformanceCounter(byte* page, ref int offset)
{
// BOOL QueryPerformanceCounter(LARGE_INTEGER* out in rcx): the emitted
// consumers only need a monotonically increasing counter, which rdtsc
// provides without leaving Win64-safe registers.
var start = (nint)(page + offset);
Emit(page, ref offset, 0x0F, 0x31); // rdtsc
Emit(page, ref offset, 0x48, 0xC1, 0xE2, 0x20); // shl rdx, 32
Emit(page, ref offset, 0x48, 0x09, 0xD0); // or rax, rdx
Emit(page, ref offset, 0x48, 0x89, 0x01); // mov [rcx], rax
Emit(page, ref offset, 0xB8, 0x01, 0x00, 0x00, 0x00); // mov eax, 1
Emit(page, ref offset, 0xC3); // ret
return start;
}
private static nint EmitSwitchToThread(byte* page, ref int offset)
{
var schedYield = ResolveLibcExport("sched_yield");
var start = (nint)(page + offset);
Emit(page, ref offset, 0x56); // push rsi
Emit(page, ref offset, 0x57); // push rdi
Emit(page, ref offset, 0x48, 0x83, 0xEC, 0x08); // sub rsp, 8
EmitMovRaxImm64(page, ref offset, schedYield); // mov rax, imm64
Emit(page, ref offset, 0xFF, 0xD0); // call rax
Emit(page, ref offset, 0x48, 0x83, 0xC4, 0x08); // add rsp, 8
Emit(page, ref offset, 0x5F); // pop rdi
Emit(page, ref offset, 0x5E); // pop rsi
Emit(page, ref offset, 0xB8, 0x01, 0x00, 0x00, 0x00); // mov eax, 1
Emit(page, ref offset, 0xC3); // ret
return start;
}
private static nint EmitSleep(byte* page, ref int offset)
{
// void Sleep(DWORD milliseconds in ecx) -> usleep(microseconds in edi).
var usleep = ResolveLibcExport("usleep");
var start = (nint)(page + offset);
Emit(page, ref offset, 0x56); // push rsi
Emit(page, ref offset, 0x57); // push rdi
Emit(page, ref offset, 0x48, 0x83, 0xEC, 0x08); // sub rsp, 8
Emit(page, ref offset, 0x89, 0xCF); // mov edi, ecx
Emit(page, ref offset, 0x81, 0xFF, 0xFF, 0x0F, 0x00, 0x00); // cmp edi, 0xFFF
Emit(page, ref offset, 0x76, 0x05); // jbe +5
Emit(page, ref offset, 0xBF, 0xFF, 0x0F, 0x00, 0x00); // mov edi, 0xFFF (cap at ~4s)
Emit(page, ref offset, 0x69, 0xFF, 0xE8, 0x03, 0x00, 0x00); // imul edi, edi, 1000
EmitMovRaxImm64(page, ref offset, usleep); // mov rax, imm64
Emit(page, ref offset, 0xFF, 0xD0); // call rax
Emit(page, ref offset, 0x48, 0x83, 0xC4, 0x08); // add rsp, 8
Emit(page, ref offset, 0x5F); // pop rdi
Emit(page, ref offset, 0x5E); // pop rsi
Emit(page, ref offset, 0xC3); // ret
return start;
}
private static nint EmitWaitForSingleObject(byte* page, ref int offset)
{
// DWORD WaitForSingleObject(worker event in rcx, timeout in edx): the
// worker loop only ever waits forever, so the timeout is ignored.
// macOS waits on a dispatch semaphore (needs DISPATCH_TIME_FOREVER in
// rsi), Linux on a sem_t; both retry until the wait succeeds (EINTR).
var wait = ResolveLibcExport(
OperatingSystem.IsMacOS() ? "dispatch_semaphore_wait" : "sem_wait");
var start = (nint)(page + offset);
Emit(page, ref offset, 0x56); // push rsi
Emit(page, ref offset, 0x57); // push rdi
Emit(page, ref offset, 0x53); // push rbx
Emit(page, ref offset, 0x48, 0x89, 0xCB); // mov rbx, rcx
var retry = offset;
Emit(page, ref offset, 0x48, 0x89, 0xDF); // mov rdi, rbx
if (OperatingSystem.IsMacOS())
{
Emit(page, ref offset, 0x48, 0xC7, 0xC6, 0xFF, 0xFF, 0xFF, 0xFF); // mov rsi, DISPATCH_TIME_FOREVER
}
EmitMovRaxImm64(page, ref offset, wait); // mov rax, imm64
Emit(page, ref offset, 0xFF, 0xD0); // call rax
Emit(page, ref offset, 0x85, 0xC0); // test eax, eax
Emit(page, ref offset, 0x75, unchecked((byte)(retry - (offset + 2)))); // jnz retry
Emit(page, ref offset, 0x31, 0xC0); // xor eax, eax (WAIT_OBJECT_0)
Emit(page, ref offset, 0x5B); // pop rbx
Emit(page, ref offset, 0x5F); // pop rdi
Emit(page, ref offset, 0x5E); // pop rsi
Emit(page, ref offset, 0xC3); // ret
return start;
}
private static nint EmitSetEvent(byte* page, ref int offset)
{
// BOOL SetEvent(worker event in rcx) -> dispatch_semaphore_signal /
// sem_post.
var signal = ResolveLibcExport(
OperatingSystem.IsMacOS() ? "dispatch_semaphore_signal" : "sem_post");
var start = (nint)(page + offset);
Emit(page, ref offset, 0x56); // push rsi
Emit(page, ref offset, 0x57); // push rdi
Emit(page, ref offset, 0x48, 0x83, 0xEC, 0x08); // sub rsp, 8
Emit(page, ref offset, 0x48, 0x89, 0xCF); // mov rdi, rcx
EmitMovRaxImm64(page, ref offset, signal); // mov rax, imm64
Emit(page, ref offset, 0xFF, 0xD0); // call rax
Emit(page, ref offset, 0x48, 0x83, 0xC4, 0x08); // add rsp, 8
Emit(page, ref offset, 0x5F); // pop rdi
Emit(page, ref offset, 0x5E); // pop rsi
Emit(page, ref offset, 0xB8, 0x01, 0x00, 0x00, 0x00); // mov eax, 1
Emit(page, ref offset, 0xC3); // ret
return start;
}
private static nint EmitExitThread(byte* page, ref int offset)
{
// void ExitThread(code in ecx) -> pthread_exit(NULL); never returns,
// so no registers need preserving. pthread_exit runs the thread's TSD
// destructors, which detaches the CLR if the thread lazily attached.
var pthreadExit = ResolveLibcExport("pthread_exit");
var start = (nint)(page + offset);
Emit(page, ref offset, 0x48, 0x83, 0xEC, 0x08); // sub rsp, 8
Emit(page, ref offset, 0x31, 0xFF); // xor edi, edi
EmitMovRaxImm64(page, ref offset, pthreadExit); // mov rax, imm64
Emit(page, ref offset, 0xFF, 0xD0); // call rax
Emit(page, ref offset, 0xCC); // int3 (never returns)
return start;
}
private static nint ResolveLibcExport(string name)
{
var libc = NativeLibrary.Load(OperatingSystem.IsMacOS() ? "libSystem.dylib" : "libc.so.6");
return NativeLibrary.GetExport(libc, name);
}
private static void Emit(byte* page, ref int offset, params byte[] bytes)
{
foreach (var value in bytes)
{
page[offset++] = value;
}
}
private static void EmitMovRaxImm64(byte* page, ref int offset, nint value)
{
Emit(page, ref offset, 0x48, 0xB8);
*(long*)(page + offset) = value;
offset += sizeof(long);
}
[DllImport("libc", EntryPoint = "pthread_key_create", SetLastError = true)]
private static extern int pthread_key_create_mac(nuint* key, nint destructor);
[DllImport("libc", EntryPoint = "pthread_key_create", SetLastError = true)]
private static extern int pthread_key_create_linux(uint* key, nint destructor);
[DllImport("libc", EntryPoint = "pthread_key_delete")]
private static extern int pthread_key_delete_mac(nuint key);
[DllImport("libc", EntryPoint = "pthread_key_delete")]
private static extern int pthread_key_delete_linux(uint key);
[DllImport("libc", EntryPoint = "pthread_setspecific")]
private static extern int pthread_setspecific_mac(nuint key, nint value);
[DllImport("libc", EntryPoint = "pthread_setspecific")]
private static extern int pthread_setspecific_linux(uint key, nint value);
[DllImport("libc", EntryPoint = "pthread_getspecific")]
private static extern nint pthread_getspecific_mac(nuint key);
[DllImport("libc", EntryPoint = "pthread_getspecific")]
private static extern nint pthread_getspecific_linux(uint key);
[DllImport("libc")]
private static extern int pthread_threadid_np(nint thread, ulong* threadId);
[DllImport("libc")]
private static extern int gettid();
[DllImport("libc")]
private static extern int pthread_attr_init(byte* attr);
[DllImport("libc")]
private static extern int pthread_attr_destroy(byte* attr);
[DllImport("libc")]
private static extern int pthread_attr_setstacksize(byte* attr, nuint stackSize);
[DllImport("libc")]
private static extern int pthread_create(nint* thread, byte* attr, nint startRoutine, nint arg);
[DllImport("libc")]
private static extern int pthread_join(nint thread, nint* returnValue);
[DllImport("libc")]
private static extern int pthread_detach(nint thread);
[DllImport("libc")]
private static extern int pthread_kill(nint thread, int signal);
// macOS: dispatch semaphores back the worker events (unnamed sem_init is
// unsupported on Darwin). libSystem reexports libdispatch, so "libc"
// resolves these like the pthread imports above.
[DllImport("libc")]
private static extern nint dispatch_semaphore_create(long value);
[DllImport("libc")]
private static extern nint dispatch_semaphore_signal(nint semaphore);
[DllImport("libc")]
private static extern nint dispatch_semaphore_wait(nint semaphore, ulong timeout);
[DllImport("libc")]
private static extern ulong dispatch_time(ulong when, long deltaNanoseconds);
[DllImport("libc")]
private static extern void dispatch_release(nint handle);
// Linux: unnamed POSIX semaphores.
[DllImport("libc")]
private static extern int sem_init(nint semaphore, int shared, uint value);
[DllImport("libc")]
private static extern int sem_post(nint semaphore);
[DllImport("libc")]
private static extern int sem_wait(nint semaphore);
[DllImport("libc")]
private static extern int sem_trywait(nint semaphore);
[DllImport("libc")]
private static extern int sem_destroy(nint semaphore);
}
@@ -0,0 +1,19 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
namespace SharpEmu.HLE.Host.Posix;
internal sealed class PosixHostSymbolResolver : IHostSymbolResolver
{
public nint GetAddress(HostRuntimeFunction function) => function switch
{
HostRuntimeFunction.TlsGetValue => PosixHostStubs.TlsGetValueStubAddress,
HostRuntimeFunction.QueryPerformanceCounter => PosixHostStubs.QueryPerformanceCounterStubAddress,
HostRuntimeFunction.SwitchToThread => PosixHostStubs.SwitchToThreadStubAddress,
HostRuntimeFunction.Sleep => PosixHostStubs.SleepStubAddress,
HostRuntimeFunction.WaitForSingleObject => PosixHostStubs.WaitForSingleObjectStubAddress,
HostRuntimeFunction.SetEvent => PosixHostStubs.SetEventStubAddress,
HostRuntimeFunction.ExitThread => PosixHostStubs.ExitThreadStubAddress,
_ => throw new ArgumentOutOfRangeException(nameof(function), function, null),
};
}
@@ -0,0 +1,57 @@
// Copyright (C) 2026 SharpEmu Emulator Project
// SPDX-License-Identifier: GPL-2.0-or-later
namespace SharpEmu.HLE.Host.Posix;
internal sealed class PosixHostThreading : IHostThreading
{
public uint AllocateTlsSlot() => PosixHostStubs.TlsAlloc();
public bool FreeTlsSlot(uint slot) => PosixHostStubs.TlsFree(slot);
public bool SetTlsValue(uint slot, nint value) => PosixHostStubs.TlsSetValue(slot, value);
public nint GetTlsValue(uint slot) => PosixHostStubs.TlsGetValue(slot);
public uint CurrentThreadId => PosixHostStubs.GetCurrentThreadId();
public void RequestTimerResolution()
{
// POSIX sleep primitives are already high-resolution; there is no
// timeBeginPeriod equivalent to request.
}
// Thread affinity is advisory on POSIX hosts (macOS offers no
// pthread-level affinity API); callers treat false as "not applied".
public bool TrySetCurrentThreadAffinity(nuint affinityMask)
{
_ = affinityMask;
return false;
}
public nint CreateNativeThread(
nint entry,
nint parameter,
nuint stackReserveBytes,
out uint threadId)
{
return PosixHostStubs.CreateWorkerThread(entry, parameter, stackReserveBytes, out threadId);
}
public bool WaitForThreadExit(nint threadHandle, uint timeoutMilliseconds)
{
return PosixHostStubs.WaitForWorkerThreadExit(threadHandle, timeoutMilliseconds);
}
public void CloseThreadHandle(nint threadHandle)
{
PosixHostStubs.CloseWorkerThreadHandle(threadHandle);
}
public bool TryCaptureThreadRegisters(uint threadId, out HostCapturedRegisters registers)
{
_ = threadId;
registers = default;
return false;
}
}