* [Tests] Add SharpEmu.Libs.Tests project
Introduce an xunit project for the HLE libs with a minimal ICpuMemory fake,
so library-level exports and helpers can be exercised without a live guest.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* [Ampr] Disambiguate pak size-collisions by read locality
PakDirectoryTracker resolves a sequential AMPR read (offset -1) back to an
absolute pak offset by matching the requested byte count against the PACK
directory. When several files share that byte count it took the first
unconsumed match in directory order, which mis-resolves out-of-order reads:
progs/h_ogre.mdl and bots/navigation/death32c.nav are both 0x3A34 bytes, and
death32c.nav sits earlier in the directory and is never read during Quake's
intro demo, so requesting h_ogre.mdl returned the nav file's bytes. The engine
then parsed "NAV2" as a brush model, failed the version check and aborted.
Pick the unconsumed same-size entry nearest the running read cursor instead.
id archives cluster related assets and the guest streams them with locality,
so this lands on the intended file; contiguous same-size runs (the
gfx/weapons/ww_*.lmp icons) still resolve in packed order.
Verified against a Quake dump: the abort is gone, h_ogre.mdl reads correctly,
and the intro demo reaches its main loop and renders instead of dying at the
error dialog.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* [Json] Implement sce::Json::Value and Json::String construct/set/destroy
libSceJson previously only had the Initializer/MemAllocator setup path.
The Value and String classes themselves were entirely absent, so a
Prospero title that builds a JSON tree (Quake PPSA01880 does, to shape
a web-API request) hit unresolved imports and faulted on the call. The
imports it left unresolved right before its access violation are exactly
these Value ctors/setters and String ctor/dtor.
Model the Value/String payload host-side (JsonObjectHeap), keyed by the
guest `this` pointer, following the handle-shadow pattern already used
by Ngs2Exports. The guest object bytes are deliberately not written:
these objects are usually stack-allocated with an unknown real layout,
and writing a guessed layout risks smashing an adjacent stack canary
(the same hazard the AudioOut2 context-param note in this tree records).
Constructors and setters follow the Itanium ABI and return `this` in rax,
which is correct whether the real setter returns void or Value&.
Covered NIDs (complete-object C1/D1 variants, matching the observed
imports): Value(default/bool/long/ulong/double/ValueType/char*/String),
Value::~Value, Value::set(bool/long/ulong/double/ValueType/char*/String),
Value::clear, String(char*/default/copy), String::~String.
Only the payload the guest can reach through library methods is modelled;
direct guest reads of the object bytes are out of scope and would need
observed layout evidence.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* [Tests] Add SharpEmu.Libs.Tests covering the Json Value/String exports
First test project for SharpEmu.Libs (xunit), the SharpEmu.Libs.Tests
layout the maintainer already agreed to in issue #36.
- A FakeCpuMemory (single contiguous region) drives the exports at the
CpuContext level with no live guest.
- Direct-call tests: ctor/setter round-trips for bool/int/uint/double
(read from xmm0)/char*/String/ValueType, destructor cleanup, and the
graceful-degradation paths (missing String shadow and a faulting char*
pointer both fall back to the empty string instead of throwing).
- Registration test: a real ModuleManager scans SharpEmu.Libs and the
nine NIDs Quake left unresolved now resolve to the libSceJson exports
and dispatch cleanly (returns `this` in rax).
InternalsVisibleTo exposes JsonObjectHeap to the test assembly. The test
project's packages.lock.json is committed for CI locked-mode restore;
CI does not run tests yet, left as a maintainer decision.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* [Json] Add Initializer::setGlobalNullAccessCallback
Quake calls it during kexPSNWebAPI::Initialize and treats the
not-found error as fatal for the whole Np Web API bring-up. Store the
guest hook (never invoked by this HLE: shadows degrade to defaults
instead of dereferencing missing members) and return success.
Verified against the dump: the "setGlobalNullAccessCallback failed
(0x80020002)" line is gone and kexPSNWebAPI::Initialize now logs
"Np Web API Initialized"; the next blockers are sceNpAuthCreateRequest
and sceUserServiceInitialize ordering, outside libSceJson.
Also pins both Json test classes to one xunit collection: they share
JsonObjectHeap statics and parallel class execution raced ResetForTests
against a running test.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* Add SaveData transaction and NP UDS layout HLE stubs
Wire Prepare, Commit, and Umount2 for implicit save transactions,
unregister guest mounts on Umount2, and add NP UDS CreateEvent,
DestroyEvent, and EventPropertyObjectSetString for layout-load imports.
* Add NP UDS SetArray and PostEvent layout HLE stubs
Add sceNpUniversalDataSystemEventPropertyObjectSetArray and
sceNpUniversalDataSystemPostEvent for layout-load imports on PPSA02929.
Move KMcEa+rHsIo from libKernel MapMemory mislabel to sceAvPlayerAddSource.
Align WV1GwM32NgY ExportName with sceNpWebApi2PushEventCreateHandle. Behavior unchanged.
Rework the sceMsgDialog and sceSaveDataDialog HLE state machines so the full
Initialize -> Open -> poll -> GetResult -> Close/Terminate lifecycle honors the
common-dialog contract, and add the three missing sceMsgDialogProgressBar* exports.
- Fix an unreachable close path: sceSaveDataDialogClose already did a
RUNNING -> FINISHED compare-exchange, but Open jumped straight to FINISHED, so
RUNNING never existed and Close could only return NOT_RUNNING. Open now enters
RUNNING and the first status poll advances it to FINISHED. Same model applied to
sceMsgDialog.
- Return the real SCE_COMMON_DIALOG_ERROR_* codes (0x80B8xxxx) from sceMsgDialog*
instead of emulator-internal result codes, with the missing argument/state guards
(ARG_NULL, NOT_INITIALIZED, BUSY, NOT_FINISHED, NOT_RUNNING).
- GetResult reports buttonId = 1 (affirmative) instead of 0, the invalid sentinel a
yes/no prompt could mis-branch on.
- Add sceMsgDialogProgressBarSetValue, sceMsgDialogProgressBarInc and
sceMsgDialogProgressBarSetMsg (NIDs wTpfglkmv34, Gc5k1qcK4fs, 6H-71OdrpXM), gated
on the service being initialized.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* [AGC] Complete gfx10 v_cmpx_f32 decode and emit ordered/unordered float compares
Add the missing v_cmpx_*_f32 VOPC decode entries (0x17-0x1C, 0x1F) and
emission for the ordered/unordered predicates: nlg maps to OpFUnordEqual,
while o/u are lowered from OpIsNan (unordered = isnan(a) || isnan(b),
ordered = !unordered) because SPIR-V's OpOrdered/OpUnordered require the
Kernel capability and are invalid in Vulkan shader modules.
Opcode numbers cross-checked against LLVM's llvm-mc regression tests
(llvm/test/MC/AMDGPU/gfx10_asm_vopc.s, gfx10_asm_vopcx.s); emitted
lowering validated with spirv-val --target-env vulkan1.1.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* [AGC] Write VCC only for non-X vector compares
On gfx10 the VCmpx encodings have no sdst and define EXEC only, so the
unconditional VCC store clobbered VCC on every VCmpx. Move the VCC store
to the non-X path; EXEC keeps the existing old-EXEC & condition update.
Addresses review feedback on #122.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: tensorcrush <tensorcrush@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Implement six missing libc string/memory search and concatenation
routines in the kernel compat layer. Titles frequently call these
during startup string handling (path parsing, config lookups, format
string assembly), and without them the loader currently falls through
to unresolved-import handling.
The implementations follow the existing byte-at-a-time compat helpers
(TryReadCompat/TryWriteCompat) already used by strcpy/strncpy/memcmp,
matching native semantics: strchr/strrchr scan through and including
the terminator, memchr is bounded strictly by count, strcat/strncat
overwrite the destination terminator and re-terminate, and strstr
returns the haystack pointer for an empty needle. NIDs are the
libSceLibcInternal/libc symbol hashes for each name.
* Expand Aerolib catalog from nids.csv and wire socket/net NID handlers
Load authoritative NID pairs from scripts/nids.csv with ps5_names fallback.
Replace mislabeled kernel zero stubs with socket/connect/bind/getsockname HLE
and sceNet byte-order exports backed by the CSV symbol names.
* Add inet_pton, htons, and bzero kernel compat with CSV NIDs
Wire libc network helpers using authoritative NID names from nids.csv
instead of synthetic Gst* exports used on the crt-loader branch.
* Fix REUSE annotation for scripts/nids.csv
* Drop bundled nids.csv; extend ps5_names and regenerate Aerolib
Remove scripts/nids.csv from the repository and fold csv-only symbol names
into scripts/ps5_names.txt so Aerolib keeps the full catalog via name2nid.
Two follow-ups to upstream #102's condition-variable changes, as analyzed
in upstream issue #113:
- The pending-signal consume path reacquired the guest mutex while still
holding the condition state lock, inverting lock order against
cond-signal (mutex -> SyncRoot) and deadlocking both threads. Leave the
condition lock before relocking, matching the normal wake path.
This unfroze Dreaming Sarah (PPSA02929) at its title screen.
- pthread_cond_timedwait's third argument is a pointer to an absolute
CLOCK_REALTIME timespec, not a relative microsecond count; the guest
address was being truncated into a duration, yielding arbitrary
timeouts. Read the timespec and convert to a relative wait.
scePthreadCondTimedwait keeps its separate relative-time ABI.
* [AGC] Decode gfx10 SOPP hint instructions
s_clause (0x21), s_waitcnt_depctr (0x23), s_round_mode (0x24) and s_denorm_mode (0x25) were missing from the SOPP decode table, so any shader containing one of these scheduling/mode hints failed to decode entirely with unknown-sopp. No emitter changes are needed: non-branch SOPP instructions are already emitted as no-ops. Opcodes verified against LLVM SOPInstructions.td (SOPP_Real_32_gfx10); decode and end-to-end SPIR-V compilation verified with a synthetic program containing all four hints.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* [AGC] Narrow SOPP additions to scheduling hints only
Per review: s_round_mode (0x24) and s_denorm_mode (0x25) write the
shader floating-point MODE state, and the emitter's blanket SOPP no-op
would have silently ignored their simm16 payloads, trading a loud
decode failure for a potential floating-point semantics mismatch. They
are removed and keep failing decode explicitly until their semantics
are modeled or conservatively validated.
s_clause (0x21) and s_waitcnt_depctr (0x23) remain: they are pure
scheduler/dependency hints with no value semantics.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* [HLE] Fix guest-thread sync and boot for Unreal Engine titles
Silent Hill: The Short Message (and other UE titles) now boot the full
engine thread graph instead of hanging early. Four related fixes:
- pthread cond/mutex semantics: retain a signal raised with no waiter as
pending, and key block/wake on the state's identity rather than a
resolved address that could differ between lock and unlock. This ends
the ~1.5M-call cond_wait busy-spin.
- Warm HLE type initializers and force-JIT their methods on a host thread
at Freeze(). A .cctor or first-time JIT running on a guest thread's
hijacked stack fail-fasts the CLR as "Invalid Program: attempted to
call a UnmanagedCallersOnly method from managed code".
- Guest thread scheduling: pump after a wake so a readied thread actually
runs, add a dispatcher thread for when every guest thread is parked,
and make the pump-depth guard an atomic CAS.
- Route mutex/rwlock lock/unlock off the non-blocking leaf-import fast
path so a contended lock can deschedule its guest thread.
Ported from the unreal-boot-fixes branch.
* [HLE] Keep mutex/rwlock unlock on the leaf-import fast path
The previous change routed all mutex/rwlock lock and unlock NIDs off the
leaf fast path so a contended lock could deschedule its guest thread. But
unlock never blocks, and taking it off the fast path made it slow enough
that Demon's Souls' job workers livelocked in a guest spinlock (millions
of mutex_unlock calls, no import progress, main thread stuck in
sceKernelWaitEventFlag).
Only *lock* needs to leave the leaf path. Restore the four unlock NIDs
(mutex + rwlock) so guest spinlocks stay cheap, while lock/rd/wrlock
remain off it for the blocking case Silent Hill needs.
* [HLE] Gate pthread_mutex_lock guest-thread blocking (fixes Demon's Souls)
Re-enabling cooperative deschedule on a contended pthread_mutex_lock
regressed Demon's Souls: its job workers run on libSceFiber, and blocking
a guest thread mid-fiber left sceFiberSwitch returning ESRCH followed by
a null fiber-context deref (0xC0000005). Bisect confirmed the pthread
change as the cause; the game reaches the same point as before it once
the block is skipped.
Gate the block behind SHARPEMU_MUTEX_LOCK_BLOCKING (off by default) so
contended locks fall through to the synchronous host-thread wait. The
rest of the pthread fixes (cond_wait pending signals, identity wake keys)
are unaffected.
Two gaps around the VOP3 signed multiplies caused whole-shader SPIR-V
compilation failures:
- v_mul_lo_i32 (0x16B) decoded correctly but had no emission case, so
any shader containing it failed with "unsupported vector opcode
VMulLoI32". Its low 32 result bits are identical to the unsigned
multiply in two's complement, so it now shares the v_mul_lo_u32 IMul
case.
- v_mul_hi_i32 (0x16C) was missing from the VOP3 decode table entirely
and decoded as an opaque Vop3Raw16C, which also fails at emission.
It is now decoded and emitted by sign-extending both operands to
64 bits, multiplying, and taking the upper 32 bits of the product,
mirroring the existing v_mul_hi_u32 pattern.
Opcode numbers verified against LLVM's AMDGPU backend
(VOP3Instructions.td): V_MUL_LO_U32 gfx10 = 0x169, V_MUL_HI_U32 =
0x16a, V_MUL_LO_I32 = 0x16b, V_MUL_HI_I32 = 0x16c. Behavior verified
by decoding and fully compiling a synthetic program containing all
four multiplies: previously the 0x16C word decoded as Vop3Raw16C and
compilation failed at the v_mul_lo_i32 instruction; now all four
decode by name and the program compiles to SPIR-V.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
VOP2 opcode 0x2B was mapped to v_ldexp_f32, which is its gfx6/gfx7
assignment. On gfx10-class hardware 0x2B is v_fmac_f32, so any shader
using it silently computed ldexp(a, b) instead of dst += a * b.
v_ldexp_f32 on gfx10 only exists as VOP3 0x362, which the VOP3 table
already maps correctly.
Also add the remaining members of the fmac family:
- v_fmamk_f32 (0x2C) and v_fmaak_f32 (0x2D), including their mandatory
literal dword in instruction sizing and operand construction, reusing
the existing v_madmk/v_madak handling.
- The VOP3-encoded form of v_fmac_f32 (0x12B), emitted when source
modifiers are present.
SPIR-V emission reuses the existing v_mac_f32 body (fma with the
destination register as addend) and the v_mad/v_fma case group.
Opcode assignments verified against LLVM's AMDGPU backend
(VOP2Instructions.td): V_FMAC_F32 gfx10 = 0x02b, V_FMAMK_F32 = 0x02c,
V_FMAAK_F32 = 0x02d; V_LDEXP_F32 is 0x02b only on gfx6/gfx7 and is
VOP3-only 0x362 on gfx10. Decode verified by feeding hand-assembled
gfx1013 words through Gen5ShaderTranslator: 0x560A0501 previously
decoded as VLdexpF32 and a v_fmamk_f32 program failed with
unknown-vop2 op=0x2C; both now decode correctly, and VOP3 0x362 still
decodes as VLdexpF32.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
SelectPhysicalDevice took the first device exposing a graphics+present
queue. On hybrid-graphics laptops that is the integrated GPU, so the
discrete card went unused - and AMD's integrated driver access-violates
inside vkCreateGraphicsPipelines while compiling some translated guest
shaders, killing the process. The CLR surfaced that native AV as
"Invalid Program: attempted to call a UnmanagedCallersOnly method from
managed code", which made it look like a CPU/threading fault.
Score the candidates instead: NVIDIA parts win, other discrete GPUs come
next, and an integrated GPU is only chosen when nothing else can present.
SHARPEMU_VK_DEVICE=<substring> pins a specific adapter, and the selected
device is logged.
* fix(vfs): defer host cursor commit in getdirentries
Follow-up VFS hardening task applying the same guest-writes-first discipline established in the time subsystem to directory enumeration.
Deferred Commit in KernelGetdirentriesCore:
- Reordered guest output so the 512-byte dirent buffer is written first via TryWriteCompat, basep is updated second (when non-null) via TryWriteUInt64Compat, and directory.NextIndex is advanced only after both guest writes succeed.
- Removed the early basep write at method entry that could mutate guest memory before buffer validation and advance the host cursor before a successful dirent delivery, causing permanent entry loss on MEMORY_FAULT at bufferAddress.
- EOF handling: when currentIndex >= Entries.Length, write basep with the final offset and return 0 without mutating NextIndex, matching FreeBSD getdirentries(2) semantics and preventing infinite retry loops.
KernelGetdents path: basePointerAddress is passed as 0, so the transaction collapses to buffer write then host cursor advance with no basep side effect.
Out of scope: coalesced {id, size} writes in sceKernelAprResolveFilepathsToIdsAndFileSizes; NetCtl connected-state stubs.
Files: KernelMemoryCompatExports.cs
* fix(vfs): resolve-first bulk commit in APR filepath resolution
Refactored sceKernelAprResolveFilepathsToIdsAndFileSizes to stop writing ids and sizes into guest memory one element at a time.
- Removed the uint.MaxValue placeholder write at the start of each loop iteration.
- Path resolution and file size lookup now fill host-side buffers first; on EFAULT or NOT_FOUND the guest ids/sizes arrays are left untouched.
- ids and sizes are packed into contiguous byte buffers and written with one TryWriteCompat call per output array instead of separate TryWriteUInt32Compat / TryWriteUInt64Compat per index.
- AmprFileRegistry.Register is called only after guest writes succeed.
- AmprFileRegistry.ComputeFileId is internal so ids can be computed without registering paths during the resolve loop.
Files: KernelMemoryCompatExports.cs, AmprFileRegistry.cs
Treat the exact untextured transparent-black premultiplied fill used by Chowdren as an overwrite. This prevents Dreaming Sarah fog and vignette render targets from accumulating across frames; SHARPEMU_DISABLE_TRANSPARENT_FILL_CLEAR=1 restores prior behavior.
Port the libSceDiscMap stubs from Kyty (InoriRus/Kyty, MIT) into the
SysAbiExport model. Disc-installed titles probe these NIDs on most file
accesses to decide whether a read must be redirected to the disc drive;
answering that every request is already resident on internal storage
keeps I/O on the regular file system path instead of failing with
unresolved-import errors.
- sceDiscMapIsRequestOnHDD (lbQKqsERhtE): validates args, writes 1 to
the result pointer, returns 0
- fJgP+wqifno / ioKMruft1ek: zero-fill the three output pointers,
return 0 (names not present in ps5_names.txt; kept as descriptive
Unknown exports like the existing sceKernelUnknown* convention)
- DISC_MAP_ERROR_INVALID_ARGUMENT (0x81100001) on null pointers,
matching the documented libSceDiscMap error range
- optional tracing via SHARPEMU_LOG_DISCMAP=1
Co-authored-by: j92580498-max <252151737+j92580498-max@users.noreply.github.com>
* [agc] WAIT_REG_MEM suspend/resume, draw packet fixes, new HLE exports, debug cleanup
Rebased onto upstream 79a7437 (par274/sharpemu, rewritten history).
- GpuWaitRegistry: DCBs suspended on unsatisfied WAIT_REG_MEM are re-polled
against guest memory on every submit; fixed 64-bit and standard packet parse
offsets, apply the mask, treat PM4 compare function 0 as "always".
- TryReadSubmittedDrawCount: accept the 5-dword ItDrawIndex2 form emitted by
DcbDrawIndex (count at +4); menu draws were silently discarded before.
- sceAgcDriverSubmitMultiDcbs: reversed ABI (rdi=address array, rsi=dword
sizes, rdx=count).
- VideoOut: vblank events, sceVideoOutGetFlipStatus, buffers registered via
sceVideoOutRegisterBuffers are valid flip targets.
- New HLE: libc stdio (fopen/fread/fseek/ftell/fclose/fgets), Dinkumware
_Getpctype ctype table, NpTrophy2 stubs, AMPR PAK sequential-read tracker,
MsgDialog lifecycle, NGS2 alt NIDs + dummy vtable for handle objects,
guarded memset intrinsic, abort()/strcasecmp null-arg recovery.
- Removed investigation-only code (INT3 breakpoints, qfont/mcpp dumps,
error-candidate printf traces, unconditional debug logs).
First rendered frame: Quake (PPSA01880) presents a 1920x1080 guest frame.
* Implemented a guarded native intrinsic (rep movsb) in DirectExecutionBackend to bypass HLE dispatch overhead, while preserving memory safety checks.
* [hle] clock_gettime clock ids, AudioOut2 canary fix, NID rebinds, new offline stubs
- clock_gettime (lLMT9vJAck0): support CLOCK_SECOND and the *_PRECISE/*_FAST
variants instead of returning EINVAL, which games treated as fatal and
retried in a tight loop.
- AudioOut2: context param writes shrunk to the guest-observed layout (the
old 0x80-byte reset smashed the stack canary at +0x60 and killed audio
init); ContextQueryMemory writes the single u64 the caller expects.
- NGS2: dropped wrong alt-NID aliases (they hash to sceImeUpdate,
sceMouseRead, sceSystemGestureUpdateAllTouchRecognizer - now bound in
their real libraries); added sceNgs2PanInit; fixed VoiceGetState NIDs.
- New verified stubs: sceUltInitialize, sceNpUniversalDataSystemDestroyHandle,
sceNpGetOnlineId, sceNpGetNpReachabilityState, sceImeKeyboardOpen,
sceImeKeyboardGetResourceId, sceMouseOpen, sceKernelAprGetFileSize.
- Import gateway: unwind guest workers at dispatch during backend teardown;
env-gated SHARPEMU_LOG_THREAD_MODE tracing.
* [cpu] Isolate guest execution on native worker threads
Guest entry stubs no longer run above CLR-managed frames: each run is handed
to a pooled raw OS thread whose loop is emitted native code. While guest code
executes there is not a single managed frame on the thread and it stays in
preemptive GC mode, so the GC never walks a frame chain interleaved with
guest stubs that carry no CLR unwind info (the ReversePInvokeBadTransition /
UnmanagedCallersOnly FailFast class of crashes on pumped guest threads).
- NativeGuestExecutor: CreateThread + emitted run loop (WaitForSingleObject,
UnmanagedCallersOnly prologue/epilogue, entry stub call, SetEvent). The
prologue rebinds guest TLS, the host-RSP slot, thread affinity and the
Active* ambient per run, so workers carry no guest identity and pool
freely; the orchestrating managed thread parks in a preemptive wait.
- All three entry sites route through RunGuestEntryStub: guest thread
entries, blocked-continuation resumes, and the main ExecuteEntry.
- Teardown stops workers before any executable stub or TLS index they
reference is freed; a worker that will not stop leaks its loop instead of
freeing running code.
- Kill switch: SHARPEMU_DISABLE_NATIVE_GUEST_WORKERS=1 restores the inline
calli path.
Implements the sceKernelNanosleep export (NID QvsZxomvUHs) for both Gen4
and Gen5 targets. Reads the requested timespec from guest memory,
validates the pointer and tv_nsec range, sleeps for the requested
duration, and zeroes the optional remaining-time struct on completion.
Also fixes: reading rqtp as a guest pointer to a timespec (tv_sec/tv_nsec
int64 pair) instead of raw register values, and keeps the optimized
sceKernelUsleep short-sleep path untouched.
Co-authored-by: par274 <par274@users.noreply.github.com>
Comprehensive refactoring of the system time subsystem to unify clock dispatching, support precise clock extensions, and secure memory boundaries against partial state corruption.
Centralized Clock Dispatch Engine:
- Extracted shared elapsed-tick calculation and clock-routing math into a unified internal static bool ResolveClockTime() dispatch engine under KernelRuntimeCompatExports.cs.
- Moved all clock identifiers from KernelMemoryCompatExports to KernelRuntimeCompatExports as internal const int constants to eliminate cross-file duplication while preserving raw compiler switch-case layout optimizations.
- Added native alias mapping support for CLOCK_REALTIME_PRECISE (9) and CLOCK_MONOTONIC_PRECISE (11).
- Hardened the Orbis sceKernelClockGettime path by routing it through the new dispatcher, resolving a pre-existing logic flaw where any non-zero clock_id incorrectly fell back to monotonic time. Invalid IDs now properly fail with ORBIS_GEN2_ERROR_INVALID_ARGUMENT.
Coalesced Single-Transaction Memory Writes:
- Replaced consecutive isolated 8-byte scalar writes across POSIX clock_gettime, gettimeofday, and Orbis sceKernelClockGettime/sceKernelGettimeofday with safe single-transaction 16-byte stackalloc byte buffer writes via BinaryPrimitives and ctx.Memory.TryWrite. This entirely prevents partial memory state corruption on virtual page boundaries.
- Implemented a single 8-byte coalesced zero-fill transaction for the deprecated/legacy timezone buffer (timezoneAddress != 0), aligning it with standard FreeBSD stub behavior.
- Standardized POSIX failure path routines. Write faults cleanly issue TrySetErrno(ctx, Efault) while safely omitting explicit manual Rax writes, letting the import dispatcher natively sign-extend the return -1 value to 0xFFFFFFFFFFFFFFFF.
Zero-Alloc Host RDTSC Execution Stub:
- Patched CreateRdtscReader() to stream native architecture opcodes out of stack-allocated spans directly into host executable memory zones (VirtualAlloc) via unsafe { Buffer.MemoryCopy(...) }, completely removing the high-frequency .ToArray() runtime allocation overhead on the hot path.
Files: KernelRuntimeCompatExports.cs, KernelMemoryCompatExports.cs
Follow-up task to enforce coalesced guest memory writes within the gettimeofday subsystem, removing remaining partial-write risks on virtual memory page boundaries.
* sceKernelGettimeofday Hardening: Replaced consecutive isolated 8-byte scalar writes with a single 16-byte coalesced transaction buffer using stackalloc byte[16] and BinaryPrimitives. It preserves native Orbis semantics by returning ORBIS_GEN2_ERROR_MEMORY_FAULT on failure states without side-effect partial-writes.
* POSIX gettimeofday Compliance:
- Applied the identical single-transaction 16-byte write pattern for the timeval structure.
- Implemented a single 8-byte coalesced zero-fill transaction for the deprecated/legacy timezone buffer (timezoneAddress != 0) using BinaryPrimitives.WriteInt32LittleEndian, aligning it with standard FreeBSD stub behavior.
- Integrated proper TrySetErrno(ctx, Efault) tracking upon write failures. The method safely omits explicit manual Rax writes on error paths, allowing the import dispatcher to cleanly sign-extend the return -1 value to 0xFFFFFFFFFFFFFFFF.
Out of scope: Subsystem clock and timeval validation is now fully complete; no further temporal partial-write vulnerabilities remain within the core runtime memory compat layers.
Files: KernelRuntimeCompatExports.cs
Replaced the no-op stub for sceKernelGetCompiledSdkVersion with a proper runtime compliance implementation.
Runtime Validation: Added explicit NULL pointer verification for the destination buffer address (versionAddress == 0). It returns ORBIS_GEN2_ERROR_INVALID_ARGUMENT and sign-extends the target Rax register to 0xFFFFFFFF80020003, strictly mirroring the PthreadJoin error-handling pattern of this subsystem.
Target-Based SDK Fallback: Implemented deterministic fallback version routing based on ctx.TargetGeneration (0x05000000 for Gen4 and 0x09000000 for Gen5 standard Orbis layout). This ensures guest applications pass early firmware checks until native metadata extraction is implemented.
Atomic Memory Write: Secured the state write sequence via the native ctx.TryWriteUInt32 layer, correctly catching virtual memory page faults, propagating ORBIS_GEN2_ERROR_MEMORY_FAULT to Rax, and safely bypassing partial-write state corruption.
Out of scope (follow-up): Native parsing of the compiled SDK version flags directly out of the guest ELF note/metadata sections.
sceKernelWaitSema parks a guest thread on the scheduler when the count is not
yet available, but sceKernelSignalSema only incremented the count and returned:
there was no WakeBlockedThreads call anywhere in the file, so a thread blocked
in WaitSema was never woken and the game hung there. sceKernelCancelSema and
sceKernelDeleteSema left parked waiters stranded the same way.
Give each semaphore a per-handle wake key and each waiter a small record with
the count it needs and a result slot. Signal, cancel, and delete wake the
waiters through the scheduler after releasing the semaphore lock, matching the
lock order the event flag and event queue paths already use. The wake handler
runs under the scheduler gate and consumes the count under the semaphore lock,
so a waiter needing more than is available stays parked while a smaller waiter
can still proceed; the resume handler hands the recorded result back as the
guest's return value.
Cancel bumps an epoch and delete sets a flag so woken waiters return what the
kernel returns in those cases: ECANCELED (0x80020055) for a canceled wait and
the EACCES-class 0x8002000D for a deleted semaphore. Delete succeeds even with
waiters present. Only the woken waiter's own handler adjusts the waiting-thread
count, so a waiter that parks during a cancel is not double-counted, and the
create path now wakes a waiter that raced onto the handle if the handle
write-back fails instead of stranding it.
This does not change the immediate paths: an available count is still consumed
inline, and a wait with a timeout pointer still returns immediately (honoring
the timeout through the scheduler is a separate change).
Verified with a block/wake harness that drives real guest threads through the
real import trampolines: signal-after-block, signal racing the park,
multi-waiter signal, need-count gating with a smaller waiter slipping past, and
cancel and delete with parked waiters including the reported waiter count, plus
event flag and event queue regression checks. Builds clean on Windows and
Linux.
* [GUI] Added Atrac9 audio decoder and improved GUI with audio preview and controller support
* fix: package.lock.json for SharpEmu.CLI to match the other projects
* fix: packages.lock.json file to include new dependencies for GUI improvements
* rollForward: "disable"
* [cpu] Implement SysV variadic float ABI (xmm0-7 capture, float returns, printf %f)
The import trampoline spilled only xmm0 and never reloaded a return xmm0. The
guest uses the System V AMD64 ABI: variadic float args pass in xmm0..xmm7 and
float/double returns come back in xmm0. As a result variadic float args past
the first were unavailable to HLE handlers, float returns never reached the
guest, and direct printf read %f/%e/%g from GP registers instead of XMM,
printing garbage and desynchronizing every following argument.
- Trampoline: spill xmm0..xmm7 into a 0x80-byte save area below the GP argpack
(r12 stays at the argpack base) and reload the return xmm0 in the epilogue.
- Gateway: read xmm0..7 from the save area into CpuContext and write the
handler's xmm0 back. XMM is caller-saved in SysV, so restoring xmm0 on return
is safe for non-float imports too.
- RegisterPrintfArgumentSource: read float args from xmm0..7 with independent
GP/FP counters and a shared stack-overflow cursor.
Every emitted byte was decoded; a unit test confirms float args read xmm0..7
(not GP) and interleaved "%d %f %d %f" stays synchronized. Build 0/0.
* [cpu] Document the scalar-only leaf-import constraint at its registration site
- IsLeafImport: spell out the no-XMM-args / no-XMM-return invariant the fast
path relies on and what breaks if it is violated; record the 2026-07-11 audit.
- Name every previously uncommented NID in the leaf list (mutex lock/unlock,
usleep, the Ampr/Apr command-buffer block, the unknown AGC packet NID).
- IsNoBlockLeafImport: document that it is a sub-filter of IsLeafImport and
that its five extra entries currently take the full gateway path; fix the
mislabeled K-jXhbt2gn4 comment (pthread_mutex_trylock, not
scePthreadMutexTrylock, which is upoVrzMHFeE).
- Point the DispatchImport call-site note at the audited list.
Comment-only change: the comment-stripped diff is empty and the solution
builds with 0 warnings / 0 errors.
Refactored parts of the time subsystem to improve POSIX/Orbis compliance and secure guest memory boundaries.
**1. POSIX `clock_gettime` updates:**
- Added `CLOCK_REALTIME_FAST` (10) and `CLOCK_MONOTONIC_FAST` (12) support for games using FreeBSD fast clock extensions.
- Fixed `NULL` pointer handling for `timespecAddress == 0`. It now returns `-1` with `EINVAL` (22) instead of `EFAULT` to match Orbis runtime behavior.
- Invalid `clock_id` values now properly fallback to `default` -> `-1` + `EINVAL`.
**2. Memory safety & monotonic tracking:**
- Replaced dual 8-byte scalar writes in both POSIX `clock_gettime` and Orbis `sceKernelClockGettime` with a single 16-byte write via `stackalloc byte[16]` and `BinaryPrimitives`. This prevents partial memory corruption at page boundaries.
- Bad non-NULL guest addresses now fail cleanly as `EFAULT` (POSIX) or `MEMORY_FAULT` (Orbis).
- Extracted core monotonic math into `GetProcessMonotonicTime()` in `KernelRuntimeCompatExports.cs` so both clock paths share the exact same `_processStartCounter` base.
**3. Host RDTSC optimization:**
- Fixed `CreateRdtscReader()` to copy stack-allocated opcode bytes into host `VirtualAlloc` memory via `unsafe { Buffer.MemoryCopy(...) }`. This completely gets rid of the redundant `.ToArray()` allocation on the hot path.
**Out of scope:** `sceKernelGettimeofday` / POSIX `gettimeofday` partial-write hardening; stricter clock validation in `sceKernelClockGettime`.
Read guest C strings in page-bounded chunks without heap allocations.
Return false when the buffer fills without a null terminator. Route exit
and _exit through RequestProcessExit.
* Pad: native DualSense support via raw HID
Read a real DualSense (or DualSense Edge) controller directly over
Win32 HID and feed its state into scePadRead/scePadReadState, replacing
the keyboard-only input path. No new dependencies.
- Device discovery by Sony VID/PID through setupapi/hid.dll, with
hot-plug: disconnects fall back to keyboard and reconnect
automatically
- USB input report 0x01 and Bluetooth extended report 0x31 (activated
via the feature report 0x05 handshake) are both parsed
- Full mapping to SCE_PAD_BUTTON conventions: face buttons, d-pad hat,
L1/R1/L2/R2 digital bits, analog triggers, L3/R3, Options, touchpad
click, both sticks
- Controller and keyboard input merge: buttons OR together, controller
sticks win past a small deadzone, triggers take the max
* Pad: rumble and lightbar output for DualSense
Wire scePadSetVibration, scePadSetLightBar and scePadResetLightBar to
real DualSense output reports. The output payload follows the same
layout as the Linux hid-playstation driver: both rumble motors,
lightbar RGB and the player LED indicator.
- USB uses output report 0x02; Bluetooth uses the 0x31 wrapper with a
sequence tag and CRC32 (0xA2-seeded) trailer, transport detected
from the first input report
- Output goes through a dedicated device handle so writes never
contend with the blocking input read loop
- On connect the controller gets a default state (blue lightbar,
player 1 LED); rumble state resets on disconnect
Verified on hardware over USB: lightbar color cycling and both motors.
Bluetooth output is implemented per spec but not yet hardware-tested.