85 Commits
Author SHA1 Message Date
jpolo1224 33343cd153 Release 0.8 2026-08-15 01:26:54 -04:00
jpolo1224 6cd7866553 i18n: Brazilian Portuguese corrections
From johnpetersa19. Fixes renderer.shaderChain.pass, which read "passar" -- the
verb "to pass" rather than a rendering pass -- and its plural, translates a
label left in English, and adds the packages.* strings added after the original
translation. Also drops five strings that were stored truncated mid-sentence;
English is better than half a sentence.
2026-08-15 01:26:54 -04:00
jpolo1224 6a5ad70ec7 Updater: pick the release asset that matches this build
A release now carries four APKs rather than one, so the updater has to choose
the asset built for the device it is running on instead of taking the first it
finds. Matches on the variant suffix in the asset name and falls through to the
next release rather than giving up when one has no usable asset.
2026-08-15 01:26:54 -04:00
jpolo1224 c990f34d5f UI: frame generation controls, and route the setting to the core at all
Adds the import row for Lossless.dll, the multiplier, Performance shaders and
Motion detail, plus the strings for all of it.

Frame Generation was not reaching the emulator. Rpcs3Bridge.setSetting is a
translation table keyed by (section, key) and anything absent is silently
dropped, so the toggle looked like it worked and did nothing. Enums also have
to cross as NAMES rather than indices -- sending "1" would have been wrong even
with the entry present. Found by an unconditional probe in the present path,
after being wrong about the cause twice; the probe printed mode=0 while the UI
held 1, which was the whole answer.

Performance shaders default ON. It selects framegen's 3.1p shader family
instead of 3.1, which is materially cheaper, and on a mobile GPU the
full-quality path costs more than the frames it buys. Both families are
extracted from the user's DLL already, so this switches between shaders that
are both sitting in the cache.

Motion detail is the optical-flow resolution, stored as a percentage rather
than upstream's divisor so the slider reads the right way round. Both take
effect when frame generation next starts, since the shader family and the flow
scale are baked into framegen's device and pipelines at initialize; the
descriptions say so.

The description also warns about the two things testers will otherwise report
as bugs: on-screen text shimmers because the overlay and the game's own menus
are interpolated along with everything else, and toggling mid-game pauses for
a few seconds while a second device and the pipelines are built.
2026-08-15 01:26:44 -04:00
jpolo1224 bbaebe47a4 UI: make the OSD colour control change the OSD, and let the position move in game
The "OSD Color" row -- on the Overlay tab and cycled from the in-game menu --
wrote `osdColor`, which is PCSX2's EmuCore/GS/OsdColor plus a
NativeApp.osdSetColor() that is an Unsupported.note() stub here. Both dead, so
the control had never done anything and the overlay sat on whatever RPCS3
defaulted to, while the real picker sat a hundred lines further down the same
tab. Both rows now drive ps3.overlayBodyColor.

The defaults were also wrong in a way that made this worse: they held RPCS3's
RGBA hex verbatim in fields that argbToRgba reads as ARGB, so every channel was
rotated one byte and #FFE138FF orange rendered as #E138FFFF. That is the pink
the overlay has always drawn in, and it applied to picked colours too, so
nothing ever matched what the user chose.

The preset row shows no selection when the colour came from the RGBA sliders,
and the in-game row reads "Custom", rather than naming a preset that is not
active. Overlay position is now cycled from the in-game menu as well -- it was
only in All Settings, unreachable at the one moment it matters, when the stats
are sitting on top of something you are trying to see.

Also carries the two frame generation settings fields, which live in the same
Ps3 settings class.
2026-08-15 01:26:29 -04:00
jpolo1224 e480c291da Emu: say more when a thread dies, and less when a game polls
The one-shot PPU state dump now follows its summary with what each PPU can
report about itself -- registers, the guest call stack, and the recent guest and
HLE/LV2 calls when PPU Calling History is on. Diagnosing the Saint Seiya stall
meant reconstructing that by hand from a log that only named the thread; cia
under the recompiler is written at block boundaries, so it names where a thread
has BEEN, not where it is, and the call history is only populated by the
interpreter.

cellSysutil's parameter query drops from warning to trace. Eternal Sonata
(BLJS10017) asks for ID_ENTER_BUTTON_ASSIGN twice every 33 ms and never stops,
which is about sixty lines a second for an entire session. Games polling this
is normal behaviour, not something to warn about, and the log volume alone is
enough to slow the emulator down.
2026-08-15 01:26:16 -04:00
jpolo1224 7d25a7086e VK: frame generation through Lossless Scaling, experimental
Interpolates frames between the ones the game draws, at x2/x3/x4. The shaders
come from the user's own Lossless.dll; nothing is bundled or downloaded.

framegen runs on its OWN VkDevice and statically links volk, which defines 655
globals named vkCreateImage, vkQueueSubmit and so on -- including all 124 our
loader declares. Linked into the core those either fail to link or, worse,
merge, and framegen's volkLoadDevice() then repoints the whole RSX renderer at
framegen's device. So it lives in libarmsx3_lsfg.so, reached only by dlopen
with RTLD_LOCAL, behind a C ABI and a version script that exports eleven
symbols and nothing else. Verify with llvm-nm --dynamic --defined-only: only
armsx3_lsfg_* may appear.

Two devices with no shared semaphore means images cross as AHardwareBuffer --
Adreno and Mali both refuse vkGetMemoryFdKHR(OPAQUE_FD) on AHB-imported memory,
so upstream's FD path does not work on this hardware. Capture costs 0.007
ms/frame CPU, measured; the cost is the synchronisation, not the copies.

Notes for anyone reading this later:

  * The shader loader's user pointer must outlive initialize(). framegen copies
    the callback into ShaderPool::source and resolves shaders lazily while
    BUILDING THE CONTEXT, so a stack local there is read back from a dead frame
    -- a segfault executing at a mapped, non-executable address.
  * The "device UUID" is not one. framegen matches (vendorID << 32) | deviceID.
    Zero matches nothing.
  * Imported shaders are cached to disk. They used to live only in the library's
    map, so every restart silently had none and generate() returned 0 before
    doing any work.
  * Capture takes the COMPOSITED swapchain image, after overlays. Capturing the
    game image put the perf overlay on real frames only, so it blinked at half
    the display rate.
  * generate() runs only on a frame the game actually drew, or the PPU/SPU
    compilation screen gets interpolated too.

The pipelined path that would take waitIdle off the critical path is present but
disabled behind k_framegen_pipelining_enabled: holding a frame back conflicts
with frame-context recycling, and at least one reclaim path has not been found.
The serialised path is what works. Frame generation costs some real framerate
and wants a steady one -- interpolating an unstable rate reads as judder -- so
it is labelled experimental in the UI.
2026-08-15 01:25:21 -04:00
jpolo1224 dbbb6fbde0 VK: use extended dynamic state to collapse pipeline permutations
Cull mode, front face, depth test/write/compare and primitive topology move out
of pipeline identity and into per-draw state where VK_EXT_extended_dynamic_state
is available. Fewer pipeline objects to compile and cache is worth a lot on
Adreno and Mali, where first-run compilation is a visible source of stutter.

Topology only collapses within its class -- triangle list/strip/fan share one
pipeline, lines share one, points stand alone. vkCmdSetPrimitiveTopology cannot
cross classes without dynamicPrimitiveTopologyUnrestricted, which comes from
extended_dynamic_state3 and is not something mobile drivers report. The class
representative is restart-aware: primitive restart on a *_LIST topology is
illegal without primitiveTopologyListRestart, so a restarting draw is
represented by the strip form or pipelines that build today start failing
validation.

Gated on the feature bit, not the extension string, and enabled at device
creation; without it the props keep their real values and the command stream is
byte-identical to before. Entry points go through the existing VKProcTable
wrangler, so vk_android_loader needs no regeneration.

pipeline_props keeps its shape: the disk cache stores it as a raw struct, so
the VALUES are normalized before it is used as a key rather than teaching
operator== about the extension. The shader cache directory becomes v1.96-eds
against v1.96 -- the suffix matters because support depends on the DEVICE, and
a driver can be swapped in through adrenotools between two runs of the same
game. Reading a normalized entry back without the extension would silently
build pipelines with culling off and depth compare NEVER.

Depth bounds, stencil, and the EDS2/EDS3 states stay static: depth bounds is
constant per device and never differentiated anything, and stencil is already
all-zero for the overwhelming majority of draws.
2026-08-15 01:25:01 -04:00
jpolo1224 d069a55acc VK: a lost surface is recoverable, not fatal
Leaving the app during a game aborted the process outright:

  Assertion Failed! Vulkan API call failed with unrecoverable error:
  Surface lost (VK_ERROR_SURFACE_LOST)   swapchain.cpp, swapchain_WSI::init()

Losing the surface is routine on Android -- the ANativeWindow is destroyed
every time the app leaves the foreground -- and the renderer already treats it
as recoverable everywhere else, setting m_surface_lost in both the acquire and
the present paths. Only swapchain init went through die_with_error.

All three surface queries in init() now return false instead of aborting, and
record which kind of failure it was. The caller needs that distinction: "the
window is minimized, retry later" and "the VkSurfaceKHR is dead" both surface
as a false return, but retrying against a dead surface queries the same dead
handle forever. Only the second recreates the surface first.

That also removes the memory corruption behind it. The fatal error killed the
RSX thread mid-operation and the Main Callbacks thread then destroyed its
objects, so tearing down ZCULL state freed a container that was still being
written -- scudo reportInvalidChunkState inside ~ZCULL_control. No fatal
teardown, no corrupted teardown.

~ZCULL_control is tightened regardless: it now drains page refs and resets prot
the way unlock_pages does, rather than freeing pages that still hold references
and leaving m_critical_reports_in_flight unbalanced -- harmless at process exit,
wrong on a restart within the same process, which is every restart here. Note
its m_pages_mutex is the only place that lock is taken; every real writer is
externally synchronized and locks nothing, so holding it must not be mistaken
for protection against a writer that is still running.
2026-08-15 01:24:47 -04:00
jpolo1224 614bf8b718 Android: re-deliver the Surface, so a missed one cannot strand the renderer
Opening a game the instant the app started left a black game area forever,
while rotating the device "fixed" it. SurfaceHolder.Callback::surfaceChanged is
a one-shot -- Android delivers it when the surface is created or resized and
never repeats -- and getNativeWindow() blocks until that single delivery
arrives, in a 100 ms sleep loop with no timeout. One missed delivery therefore
parks the RSX thread for the rest of the session. A rotation only helped
because a configuration change forces a fresh surfaceChanged.

EmulationSurface now re-delivers holder.surface on attach and on window
visibility changes. It is idempotent: the native side compares the incoming
ANativeWindow against the one it holds and no-ops on a match, so this costs
nothing when the first delivery already arrived. It has to be post()ed, since
onAttachedToWindow runs before layout and a 0x0 report is explicitly ignored.

The wait loop also logs now, every three seconds, because the failure was
otherwise completely silent: the emulator log stopped dead just after Vulkan
device creation, the perf sensor read 0.0% CPU, and nothing said why. Diagnosis
took a screenshot and dumpsys SurfaceFlinger to establish the surface existed.

Adds GSFrameBase::display_epoch, bumped when the native window is replaced. The
swapchain is rebuilt on a size mismatch and nothing else, so a replacement
window at identical dimensions was invisible; platforms that cannot swap a
window under a live swapchain keep the default and are unaffected.
2026-08-15 01:24:33 -04:00
jpolo1224 cce09dbb39 SPU: recover from a failed analysis, and stop the log floods
Three faults that showed up in tester logs, all of which made the emulator
look broken in ways the log then hid.

Eternal Sonata flooded with SPU "Invalid code" errors: when the analyser
produced no data the recompiler had an empty branch with a TODO where the
fallback belonged, so the block was neither compiled nor marked, and the same
address was retried forever. It now marks the block failed and lets the
interpreter take it -- 6320 errors in one session down to none.

The unknown-instruction and halt messages are rate-limited, per opcode and per
address rather than globally, so a repeating fault reports once instead of
every execution. One tester's log went from 600 MB to 2.0 MB; the log volume
itself had been slowing the emulator, so this is not only a readability fix.

ARM64 fault classification in Thread.cpp preferred a heuristic comparing
si_addr against the PC, which misreads a genuine data fault as an instruction
fetch. It now decodes ESR first and only falls back to the heuristic, and an
SPU halt at the 0xffdead00 sentinel is reported as a guest assertion rather
than a host segfault. BLEACH crashed here, and the misclassification gated
every recovery path behind it.
2026-08-15 01:24:19 -04:00
jpolo1224 b5a715adcf PPU: give the AArch64 register scavenger the spill slot it needs
Saint Seiya: Sanctuary Battle (BLES01421) stalled partway through PPU
compilation and booted to a black screen. The failure was in LLVM, not here:
on AArch64 the register scavenger ran out of registers under the GHC calling
convention, which pins most of the GPRs to guest state, and
AArch64FrameLowering::determineCalleeSaves returns early for GHC before it can
create the emergency spill slot the scavenger falls back on. The scavenger then
aborts, and because that takes down the whole MODULE rather than one function,
every function in it drops to the interpreter -- the boot never finishes, or
the game runs at interpreter speed with nothing in the log to explain it.

Fix creates the spill slot for GHC frames that actually need stack. 231/231
modules compile for Saint Seiya, and Sonic Unleashed's FMVs work for the same
reason. Because it is a codegen fix rather than a per-game workaround, any
title that hit this benefits.

The change itself lives in the LLVM submodule, whose remote is upstream
llvm/llvm-project, so it cannot travel in this repository. It is preserved
here as 3rdparty/llvm/armsx3-aarch64-ghc-emergency-spill.patch, applied
against the pinned submodule commit; a build without it applied will exhibit
the original stall.

Also bumps the ARM64 codegen cache version so caches produced before the fix
are not reused, and carries the PPUTranslator changes the same work needed.
2026-08-15 01:24:06 -04:00
jpolo1224 eb54f9b75a Build: four release variants, and what the new one needed
Splits the Android release into legacy / a11 / a13 / a15 so a device can take a
build matched to its CPU and OS instead of one binary suiting everything.
android/build-variants.sh drives all four from a single table of
(ndk, api, -march, apk suffix), and ConfigureCompiler.cmake takes -march per
variant rather than hardcoding one.

The legacy variant had never been compiled before: every release up to 0.7.2
was built at the gradle default of minSdk 33, so nothing had ever targeted a
lower API. Doing so turned up std::aligned_alloc, which is API 28+ -- below
that <cstdlib> does not declare it at all and the using-declaration fails to
resolve. posix_memalign is the older spelling and its result frees with plain
free(), so the rest of the header is unaffected. Kept even though legacy now
targets API 30, because it costs nothing and the next person to try a lower
floor should not rediscover it.

legacy targets armv8.1-a, which is the floor this codebase compiles at rather
than a preference: util/simd.hpp uses SQRDMLAH (v8.1 RDMA) and util/asm.hpp
has inline LSE atomics, so armv8-a does not build. Its value is cores that are
ARMv8.2 without the OPTIONAL fp16 and dotprod extensions the other three
variants require. Cortex-A53/A72/A73 class parts stay out of reach until those
two paths gain fallbacks.
2026-08-15 01:23:53 -04:00
jpolo1224 0a9fd15b57 Merge branch 'pr41' 2026-08-14 12:21:02 -04:00
Zulux91 0821bbf956 Emu: complete abandoned UE3 HD-cache install at boot (Larry: Box Office Bust)
Leisure Suit Larry: Box Office Bust (BLUS30331) copies its disc asset tree into
an on-HDD cache during a short boot window and abandons the copy when emulated
I/O is slower than a console, then crashes at "New Game" on the missing packages
(upstream RPCS3 #14402). Finish that copy once, at boot, before the guest runs.

complete_ue3_hd_cache() runs in Emulator::Load after the bdvd+hdd0 mounts and
before Run(). It is gated to a verified title-ID allowlist ({BLUS30331}):
PS3TOC.txt is a generic UE3 marker, so keying on it alone would act on other UE3
discs and build the write root from an unvalidated PARAM.SFO TITLE_ID. It parses
the disc PS3TOC.txt manifest, confines each entry textually (rejecting
traversal/drive/UNC/reserved names), copies each not-yet-complete asset
atomically via fs::pending_file, and stamps a 0-byte <file>__time sidecar to the
disc source mtime, mirroring the guest's own completeness convention.
Completeness is keyed on the sidecar AND the dest byte size, so a guest-truncated
payload is re-copied rather than skipped. On any parse/stat/space/copy failure it
returns install_failed after Kill(false), like the sibling post-ready error
exits, so the boot aborts cleanly instead of handing the guest a half-install.

For any other title the function returns after a single title-ID comparison,
before any filesystem access.

Validated on-device (Odin 3, Adreno 830): cold cache -> 800 files / 1847 MiB
copied in ~29s -> New Game reaches the Prologue, 0 access violations; 2nd boot
does no work (idempotent); a forced install_failed tears down cleanly with no
crash; Lollipop Chainsaw and Mirror's Edge boot unaffected (completer inert).
2026-08-14 10:59:34 -05:00
Zulux91 b29810d1a5 RSX: second hardening round for the semaphore wait, from re-review
A blind re-review of the previous commit (six lenses, fresh reviewers)
found real gaps in the hardening itself. Addressed here:

- The EVTSTRM gate failed open to the spin: with the event stream absent
  on a core whose armed WFE does not park, disabling the fallback
  reinstated the original full-rate spin. The paced tier now degrades to
  a 100 us scheduler sleep instead, which also keeps the timeout and
  service polls running at a bounded cadence.

- Gate the FIFO-idle wait_for_event() the same way (three reviewers
  independently flagged the contradiction between asm.hpp's new
  precondition and this ungated sibling). Without the stream it yields,
  which is that path's pre-WFE behavior.

- Non-Linux ARM64 now defaults to the previous commit's behavior instead
  of silently disabling the fallback: the false default was a regression
  against 002a9b274 on the Apple Silicon and Windows-on-ARM targets, and
  no HWCAP equivalent exists there to probe.

- The loop's snapshot is now read through the existing atomic reference
  (relaxed observe()) instead of a plain reference: the previous form was
  a formal data race whose correct codegen depended on an unrelated
  virtual call staying opaque to the optimizer.

- Guard unaligned semaphore addresses on the acquire path: exclusive
  loads fault on unaligned addresses, semaphore_release already rejects
  them, and acquire did not. Unaligned waits now use the paced tier only,
  with a warning.

- Surface the probe in the startup capability string (EVTSTRM-on/off) so
  every log records which wait shape was selected; previously the three
  possible states were indistinguishable in any output.

- Log the first-observed semaphore value in the recovery-timeout message
  as well; the previous message could not distinguish a value that
  changed during the wait from one that never moved.

- Comment corrections: the post-budget wake-on-write claim now states the
  pacing-period bound honestly; the x86 note names the yield fallback on
  CPUs without waitpkg/mwaitx; the event-stream period is stated as a
  kernel-dependent range. Note the previous commit's claim that x86 was
  unaffected was wrong: the snapshot change lets the x86 early-out fire
  where it previously compared a value against itself; the direction is
  an earlier return when the semaphore changed during the prologue.

Device check (Odin 3, ME menu, 30 s): 168.0G instructions, and the new
capability line reads EVTSTRM-on, proving the paced branch was live in
the measured run. Known residuals (ledgered, out of scope): HWCAP is a
boot-time global while the stream enable is per-CPU (migration edge);
no parking-core device has been measured; no automated test covers the
path.
2026-08-14 05:00:59 -05:00
Zulux91 5ef731c9e5 RSX: harden the semaphore event-stream fallback after adversarial review
Findings addressed (blind review, 8 lenses, see PR discussion):

- Gate the fallback on HWCAP_EVTSTRM (new utils::has_wfe_event_stream()).
  The park's wake bound is the kernel's architected timer event stream; on
  a kernel that does not enable it, a monitor-less WFE parks until the next
  unrelated interrupt. Such devices now keep the pre-existing armed-spin
  behavior instead.

- Fall through from the event-stream park to the armed one-shot instead of
  else-ing around it. On cores where the armed WFE parks, this re-arms the
  exclusive monitor every iteration, so wake-on-write is preserved even
  after the spin budget is spent; on Oryon the extra call returns
  immediately and costs nothing measurable. This also shrinks the window
  in which a written-then-overwritten semaphore value could go unobserved.

- Fix the spin's early-out: the call passed a freshly re-read value as
  old_value, which the compiler sank to immediately before the ldaxr,
  making the compare a self-comparison that never fired (verified by
  disassembly). The loop now snapshots its top-of-iteration read and
  passes that, so an already-changed value returns without waiting on
  every core class.

- Move spin_budget under ARCH_ARM64 (silences -Wunused-variable on x86).

- Log awaited and observed values in the driver-recovery timeout message,
  so a timeout caused by a transient value is distinguishable in reports.

- Rewrite the stale comments in place: spin_on_cacheline_once's event-
  stream rationale is core-class dependent (measured non-parking on
  Oryon); wait_for_event's usage rule now covers the sustained-idle
  fallback shape and names the HWCAP_EVTSTRM precondition.

Device check after hardening (Odin 3, ME menu, 30 s): 164.1G instructions
vs 179.5G for the previous commit and 402.7G pre-fix - the win holds.
2026-08-14 04:27:06 -05:00
Zulux91 002a9b274a RSX: fall back to event-stream wait when the semaphore spin does not park
The one-shot cacheline wait (ldaxr-armed WFE) used in semaphore_acquire
does not park on every core. Measured on Snapdragon 8 Elite class (Oryon,
Odin 3): WFE returns immediately while the exclusive monitor is armed
(~28.8M wakes/s in a standalone microbenchmark, vs ~20-30k/s for bare WFE
and sevl+wfe), so the acquire loop ran at ~57M iterations/s through waits
averaging 33 ms - about 99% of the RSX thread's wall time at a menu, with
each iteration also paying the driver-recovery get_system_time() check.

Keep the armed one-shot for the first 500 iterations of a wait - on cores
where it parks it keeps its instant wake-on-write, and where it does not
it acts as a short spin that still catches quick signals - then fall back
to wait_for_event(), which parks on both classes and bounds wake latency
at the architected event-stream period (~50 us measured).

Measured on device (Mirror's Edge, MT RSX on, state-verified windows):
menu instructions -55% (402.7G -> 179.5G per 30 s), played-gameplay
instructions -29% (391.6G -> 279.7G), loop iterations down ~1,400x, wait
counts/durations unchanged, 30 fps frame pacing unchanged (max frametime
34.2 ms). Note: cpu-cycles PMU counts at full clock during WFE park on
this SoC, so cycle-based profiles cannot see this change; measure with
instructions retired.
2026-08-14 03:47:48 -05:00
jpolo1224 39ca5cdab6 0.7.2: settings fixes, Oboe by default, and a working per-section Reset
Per-section Reset did nothing on most tabs. The field lists describe the tabs
as they were before the PS3 rewrite, so Reset was clearing settings the tabs no
longer show while missing most of what they do: Performance listed 22 of 47,
Graphics 45 of 57, Audio 10 of 16 -- audioRenderer, audioFormat, audioChannels
and audioCubebBackend were absent, so changing the audio backend and pressing
Reset was a no-op. Regenerated from what each tab actually writes, mapping
ps3.foo to its ps3Foo key and validating every entry against the serialiser.
Five keys also moved off Graphics because another tab owns them, which was a
cross-tab clobber waiting to happen.

Full Diagonal Range, per stick, on by default. A full diagonal was capped to
the unit circle at ~0.707 per axis, which is what a circular-gated DualShock
really sends -- but games that deadzone each axis separately then ignore
diagonals, and Oblivion's camera crawled diagonally while the cardinals were
fine. Off restores the hardware curve.

Oboe is the default audio backend on Android, with a migration for anyone still
on the old Cubeb default; a deliberate choice of another backend is kept.

Enter Button Assignment (circle/cross) is exposed. The core has always had it
and Android never showed it.

Reset all settings, in General. Per-game overrides and controller binds are
deliberately left alone -- they are invisible from that page.
2026-08-13 22:29:33 -04:00
jpolo1224 b82432c793 VK: allow native fp16 on Adreno drivers that accept it
Oblivion's water did not draw on Vulkan and did draw on OpenGL. The only
Vulkan-only shader workaround in play is the blanket disable of native float16
on every mobile GPU, which emulates it with fp32; its own comment claimed that
"renders correctly", and it does not.

The disable exists for a real failure -- Qualcomm's compiler rejected SPIR-V
containing float16_t and every pipeline came back VK_ERROR_UNKNOWN, which
presents as a black screen with working audio and a working compile overlay,
so it reads as a renderer bug rather than a shader one. That is not worth
reintroducing blind, so this is a version gate rather than a removal:

  Adreno on driver 512.676.53 or newer -> native fp16 (verified)
  older Adreno                         -> unchanged
  Mali, PowerVR, Xclipse, the rest     -> unchanged, untested either way

Found by switching the renderer to OpenGL, which isolated it to the Vulkan path
in one run after the settings-level suspects had all come back empty.
2026-08-13 22:29:17 -04:00
jpolo1224 7f54855b7d lv2/vm: fix a read-only unlink lockup, and three log floods
sys_fs_unlink handled notdir and noent but not readonly, so it fell through to
fmt::throw_exception and killed the PPU main thread inside the syscall. The
emulator then sat with nothing to run: the game froze with the CPU at 1% and
nothing in the log but a stalled RSX. On Android /app_home is the mounted ISO,
which is read-only, so any game deleting a file in its own directory hit it --
Oblivion removes warnings.txt at startup and never got past it. Returns
CELL_EROFS now, which is already what sys_fs_write and friends do. sys_fs_mkdir
and sys_fs_rmdir carried the identical block and are fixed with it.

Three log floods, all of which stall the emulator outright because writing them
is not free on Android:

- sys_fs_utime logged two warning lines per call and rides a polling loop.
  Oblivion's FileCaching thread hit it 7274 times in ten seconds on one .BSA,
  ~22k lines, and the frame loop stopped for over twenty seconds. Now trace.
- vm::lock_sudo reported a failed mlock on every mapping. Android never grants
  RLIMIT_MEMLOCK to apps, so it fails forever while advising the user to raise
  a limit they cannot raise -- 6470 lines in ten seconds here, and ~1200 in
  every other game log looked at. Reported once per session now.
- sys_mmapper's map/unmap pair, 12431 lines over the same window. Now trace.

None of them lose information: raise the channel to Trace to get them back.
2026-08-13 22:29:06 -04:00
jpolo1224 0819f1ef15 RSX: return the renderer to the 0.6 path, keeping the FIFO idle fix and ADPF
Testers consistently report the best performance on the build with the 0.6
renderer, so 0.7's graphics work goes back out. The Arkham City measurement
behind it (62.8 -> 51.2 ms) was one game on one device and did not survive
contact with a wider set of hardware.

Two files are kept from 0.7 because neither is render pass work and both are
measured wins on their own: RSXFIFO's idle spin plus WFE park, which took ~11%
of total CPU off sched_yield, and RSXThread's ADPF feed, without which the
performance-hint setting reports nothing and does nothing.

Everything else under Emu/RSX is byte-identical to 0.6. The removed work is not
lost -- it is in c4b45eee2 and can come back a piece at a time with testing
behind each one, which is how it should have gone in the first place.
2026-08-13 22:28:52 -04:00
jpolo1224 8ee20d91d5 0.7.1: remove vertex cache retention, keep the rest of the renderer
The previous commit reverted all of Emu/RSX to 0.6, which was more than the
bug required. Bisecting had already shown the render pass work was not
responsible -- reverting it alone changed nothing, while removing retention
with the render pass work in place fixed both reported games.

So only retention goes. It reused vertex cache entries across frames, and on
0.6 the attribute ring was too small for it to engage; raising the ring to
192M switched an existing path on in every game at once and handed draws
stale geometry. Back to purging every frame, as 0.6 did.

This restores what the wider revert had taken out for no reason: the render
pass reduction, the RSX FIFO idle fix, ADPF frame timing, the ZCULL and
occlusion query fixes, the swapchain and surface lifetime ports, and VRAM
budgeting.

Sonic Unleashed still does not render FMV cutscenes. That reproduces with
the 0.6 renderer too, so it is unrelated and still open.
2026-08-13 19:51:48 -04:00
jpolo1224 4797ad8a9a 0.7.1: revert the 0.7 renderer to 0.6
0.7 introduced corruption in several games that were fine on 0.6 -- flashing
and flickering in Sonic Unleashed and Dragon Ball among others. Emu/RSX is
returned to its 0.6 state in full; everything outside the renderer is kept.

The main cause was vertex cache retention. On 0.6 the attribute ring was too
small for retention to engage, so raising the ring to 192M did not add a code
path, it switched an existing one on in every game at once, and reusing stale
vertex data is what the flashing was.

Bisecting also showed the render pass work was not responsible: reverting it
alone changed nothing. It can come back, but on its own and with testing
behind it rather than as part of a batch.

Kept from 0.7: the ARM64 PPU float to integer fix, the PPU cache build
identity, the SPU checksum and block state fixes, Oboe, ADPF, and the crash
and stability ports.

Sonic Unleashed does not render FMV cutscenes. That reproduces with the 0.6
renderer as well, so it is not from any of this and is still open.
2026-08-13 19:46:13 -04:00
jpolo1224 c4b45eee27 0.7: SPU and RSX fixes, Oboe audio, and the ports from ouroboros420 and rfandango
SPU: the ARM64 block checksum folded two thirds of every block through
absolute difference, which is not injective, so adding the same value to
two words left the checksum unchanged and similar job binaries hashed
alike. Plain summation now. This is what Precise SPU Verification was
working around, and that setting is exposed properly instead of only being
reachable by hand editing the config.

SPU: a block is no longer marked permanently failed when the trampoline
rebuild fails. The compiled function was live, the state was not
recoverable for the rest of the session, and the claim could never be
retaken.

RSX: render pass churn cut in heavy scenes, roughly 113 to 85 passes per
frame. On a tile based GPU every pass boundary is a full tile store and
reload. Two Vulkan specification violations fixed, and a read/write hazard
on the render pass path.

RSX: the FIFO no longer burns a core on sched_yield while idle.

Android: ADPF is implemented rather than an inert setting, logcat no longer
allocates and makes an IPC call per line, and Silence All Logs is available
for playable titles.

Audio: Oboe backend, for the per device quirks database and stream recovery
on disconnect and route change.

Ported from ouroboros420/rpcsx: GPU Turbo, power and thermal handling, the
crash and freeze fixes, savestate and WSI surface lifetime, honest RAM VRAM
budgeting, the persistent SPU object cache design, occlusion query and RSX
fixes, frame pacing and tiler tuning.

Ported from rfandango/rpcsx: the Turnip ZCULL deadlock fix and ARM64 SPU
checksum handling.

Individual commits are credited in comments at each site.
2026-08-13 18:36:41 -04:00
jpolo1224 87ccdb8515 PPU: stop inverting float-to-int saturation on ARM64
FCTIW, FCTIWZ, FCTID and FCTIDZ carried a saturation correction that only
makes sense on x86. cvtsd2si returns 0x80000000 for any value it cannot
represent, so the result is XORed back into 0x7fffffff on overflow.

FCVTNS and FCVTZS already saturate on their own, so the same XOR turned a
correct result into its opposite: every overflowing conversion produced
INT_MIN where it should have produced INT_MAX. The mask is a no-op when
there is no overflow, so it never did anything except break that case.

Armored Core: For Answer put the player under the floor in the tutorial
because a coordinate that should have clamped high arrived clamped low.

The cache needed a build identity as well. Its key is the executable's
SHA-1 plus a settings bitset and nothing more, so the first attempt at this
fix silently reused objects compiled by the previous build and looked like
it had done nothing. Every earlier PPU codegen change had the same problem
for anyone with a warm cache.

Module loading also no longer abandons the remaining modules after one
object fails to load.
2026-08-13 18:36:26 -04:00
jpolo1224 e10f846924 App: make pause, restart and the FPS cap do what they say, and add trophies
Pause reached the core for the first time. Rpcs3Bridge.pause() set a bool and
returned, on the belief that RPCS3 has no explicit pause entry point -- Emu.Pause()
exists and _rpcsx_surfaceEvent has always called it on surface loss, which is why
backgrounding the app was the only thing that paused. Exported as _rpcsx_pause
through all four layers; resume already reached the core, so the pair was asymmetric.

Restart no longer crashes: setCustomDriver dlclose'd the previous driver handle, and
applyRendererPrefs re-applies the driver on every start, so restart unloaded the
library VMA had resolved vkGetPhysicalDeviceMemoryProperties2 out of. ~VKGSRender then
freed its heaps and UpdateVulkanBudget called into an unmapped mapping. An ICD cannot
be unloaded while anything resolved from it is reachable, so it is no longer closed.

Restart no longer returns to the library either: shutdown() set stopRequested, called
kill() and returned with the VM still live, so the run loop's finally started the
replacement and the in-flight teardown killed it -- two BootGame calls, then Unloading
ISO, by which point the restart flag was spent. shutdown() now waits (bounded) for the
core to report Stopped, and the restart is queued on vmStopControl behind it.

FPS cap applies at every value. ConfigStore recorded a persistent core override of
Video@@Frame limit=60 and Settings rewrote it on every push, both from a migration
escaping a stored 120 -- but that node is the cap control, and overrides replay last,
so presets were pinned at 60 while 20 and 45 worked through Second Frame Limit. The
Vblank Rate force stays, since Frame limit Auto resolves to it. Stale overrides are
cleared once. 90 and 120 dropped from the row: the min() in the pacer discards them.

Cover art for PKG installs: the library grid's fallback chain stopped one leg short of
the extracted ICON0.PNG while the in-game menu's did not. Both now share one chain, so
they cannot diverge again. has()/discIconFile require bytes rather than existence, and
the staging rename is checked instead of discarded.

Licences are grouped per game and collapsed instead of a flat list of content ids.

Trophies: a library-wide browser and an in-game tab for the running title, reading
TROPCONF.SFM and TROPUSR.DAT directly -- no account, no network. The in-game set is
identified from the core's own current_trophy_name (try_get, since get<> would
construct it outside emulation and hand back an empty name), falling back to TROPDIR
on disk because a game registers its context lazily. Note the entry stride there is
16 + entries_size, not entries_size.

Also: renderer.upscale.label was defined twice, so Internal Resolution was dead.
2026-08-12 16:46:01 -04:00
jpolo1224 db6ee86806 Core: survive a module LLVM cannot compile, and fill in the Android string table
PPU: a module that fails codegen no longer takes the boot with it. ppu_initialize2
called the fatal jit.add(); run_recoverable_llvm and the try_* pair already existed
in this tree but were used only by the SPU recompiler, so LLVM's fatal handler threw
on a thread with no recovery context and killed the worker. That is not one lost
module: g_progr_pdone is incremented in the compile loop's INCREMENT, so the module
the dead worker held was never accounted for, g_progr_ptotal could never reach zero,
and the boot waited on it forever. Saint Seiya: The Sanctuary (BLES01421, issue #25)
stopped at 133 of 134 on 'Cannot scavenge register without an emergency spill slot'.
Now routed through try_add on ARCH_ARM64, mirroring the SPU branch, with
ppu_initialize2 returning bool so the caller stops logging a dead module as compiled.
Losing the worker also halved the rate for everything left.

VK: a data_heap block no longer frees through an allocator that is not the current
one. Borrowed pointer, cached at construction with nothing tying it to the
allocator's lifetime; declining the free costs nothing the device teardown does not
already release.

Android: g_strings held 180 of localized_string_id's 323 entries and the callbacks
ignored their args entirely, so every string carrying a name, date, size or error
code lost it -- including CELL_SAVEDATA_LOAD, which is why the save prompt was Yes
and No over an empty message (open_msg_dialog logged msgString=""). All 322 the Qt
switch provides are present, in enum order, with QString::arg's %0 substitution
reproduced and utf8_to_u32string on the u32 path so trophy names survive. A
static_assert on the table size fails the build when upstream adds an id.

probeDiscInfo: set g_fxo up before mounting. vfs::mount lazily constructs vfs_manager
through manual_typemap::init<T>(), which writes *m_order++, and clear() nulls that
when a game stops -- so scanning a new disc image after playing anything wrote
through null. Emu.IsStopped() cannot guard it, because stopped is the cleared state.
2026-08-12 16:45:29 -04:00
jpolo1224 708582e523 Ship the ANGLE libraries from the module that actually builds
libEGL_angle.so and libGLESv2_angle.so lived in armsx3-app, which stopped being the
built module, so selecting ANGLE for the OpenGL renderer silently fell back to the
system driver with nothing in any log to contradict it. Moved into armsx3-ui beside
the core, with the jniLibs .gitignore negations that keep them tracked.

verifyAngleLibs comes with them and now runs on the release graph ahead of
mergeReleaseJniLibFolders, so packaging an APK that offers ANGLE without shipping it
fails the build. The copy left behind in armsx3-app could never have protected
anything from there, and did not even compile -- its GradleException message escaped
'$' as if the file were a template, and the quotes inside the escaped interpolation
closed the string early, so the project failed to configure. Deleted rather than
fixed, with a comment pointing at the live one.

Version to 0.6 (versionCode 10).
2026-08-12 16:45:12 -04:00
jpolo1224 8a6deab362 Touch: put the tap-to-reveal pause row back in the in-game menu
Its comment was still there, above the OSD selector, describing a control that no
longer existed -- the row was lost in the port and the setting left with no writer,
so nobody could hide the glyph or bring it back. Reported as the option missing from
the menu, which is what it was.

Goes where the comment says rather than in the touch editor toolbar, which is where
I first put it: this is a pause-button behaviour toggle and belongs with the overlay
controls it was written for.
2026-08-12 10:22:35 -04:00
jpolo1224 5b740f8921 Touch: give tap-to-reveal pause a control
The setting has existed since the pause button moved to the top right, and is seeded
once from the old show/hide pref so anyone who had the button hidden keeps it hidden.
Nothing ever wrote it afterwards. A user whose button was visible had no way to hide
it and a user migrated into hidden had no way back, which is how it was reported:
the option is not in the in-game menu.

Sits with multi-touch, gliding and floating stick in the editor toolbar, since those
are the other whole-overlay behaviour toggles and setPauseTapToReveal already
existed to be called.
2026-08-12 10:16:05 -04:00
jpolo1224 11f043b529 VK: retire completed frames on flush again, so the upload rings reclaim
The freeze-with-audio in Ratchet & Clank is the RSX thread dying in the allocator,
and the heap growth log says why. The index buffer went 16M to 64M to 128M to 192M
to 256M inside 290ms, on requests of 2K, 4K, 5K and 3K; the attrib buffer did the
same and died growing to 192M. Kilobyte allocations cannot need a quarter gigabyte.
The rings were never wrapping, they were only ever growing.

frame_context_cleanup is what returns a frame's ring memory, and check_present_status
is what calls it. I removed that call from flush_command_queue in 0.5 because the
drain poked the oldest queued frame's fence and on Adreno vkGetFenceStatus blocks
until signalled instead of answering -- 14.6ms a frame, second only to the FIFO decode
loop. The reasoning was that the flip path retires frames anyway. It does, enough to
keep presenting, but not often enough to keep the rings bounded, and nothing else
reclaims them.

Restoring it costs nothing now. poke() no longer asks with vkGetFenceStatus: it uses
vkWaitForFences with a zero timeout, which is specified to return VK_TIMEOUT without
waiting. The measurement that motivated the removal was of the old implementation, so
the speedup stays and the reclaim comes back.

Keeps the heap growth log that found this. The allocator reports only a size and a
pool number, and pool 1 covers every data_heap, so three fixes were aimed at a target
that could not be seen. One line naming the heap settled it.
2026-08-12 10:02:13 -04:00
jpolo1224 b8987b8c92 Revert "VK: grow upload heaps in steps a phone can actually place"
This reverts commit 403df1c651.
2026-08-12 09:54:00 -04:00
jpolo1224 403df1c651 VK: grow upload heaps in steps a phone can actually place
Ratchet & Clank freezes with audio still playing, which is the RSX thread dying:
'Failed to allocate 131072K of video memory (pool=1, pool total=561M, heap cap=
2048M)'. Pool 1 is VMM_ALLOCATION_POOL_SYSTEM, and 131072K is a data_heap taking
its second growth step, 64M to 128M.

The device is not out of memory. It holds 561M against a 2048M cap and cannot place
128M in one piece, which is a different failure from being full and has a different
fix. The heap grew by aligning up to 64M, so every growth demands a single
contiguous block of at least that size, and each step doubles what the allocator has
to find unbroken. A heap that fails to grow has nowhere to degrade to, so the
renderer ends there.

Android now grows in 16M steps and stops at 256M. The smaller granularity asks for a
quarter as much contiguous memory per step and lets the heap settle near the size
actually wanted rather than overshooting to the next 64M boundary. The ceiling comes
down to match: a 1GiB upload ring would exhaust the device long before it was
reached, so as written it was a limit only reachable by dying. Desktop keeps 64M and
1GiB.

Does not touch the separate pool-0 case fixed in the previous commit, where recovery
does run and the last-ditch eviction now gets a turn before the thread is killed.
2026-08-12 09:50:26 -04:00
jpolo1224 614cdd74a3 VK: evict everything before ending the RSX thread, not after
Ratchet & Clank freezes with audio still playing, which is the RSX thread dying on
VK_ERROR_OUT_OF_DEVICE_MEMORY while the rest of the process lives. Caught on an
Adreno 740: 'Failed to allocate 86016K of video memory (pool=0, pool total=472M,
heap cap=2048M)'. One 84MB request refused while we held 472MB of a 2048MB cap, so
the heap was not full -- a single large allocation could not be placed.

The allocator already retries once after asking for pressure relief, and it did.
The relief is what fell short. on_vram_exhausted refuses the hard sync whenever the
RSX is uninterruptible, and clamps the request below fatal, so the eviction that
drops everything unlocked was unreachable from here. That refusal is right while
there is still a way out: eviction touches resources the driver may still be
reading. It is wrong on the last attempt, where the alternative is not a glitch but
the renderer ending.

So the final attempt is now exempt, through a thread-local set only for the width of
that call. Everything else keeps the existing behaviour, and a caller that opted out
of recovery is not handed it here by the back door. Recovery logs at error level and
says a visual glitch is the expected outcome, since a silent recovery that costs
texture quality reads as a new bug otherwise.

Measured against this failure the eviction ran once, six microseconds before the
allocation failed, and never in the five minutes before it -- so nothing was
reclaimed while it still would have been cheap. That part is not addressed here: the
budget-based ladder cannot see mobile unified memory, where the driver reports one
large shared heap and 472MB against it never crosses a threshold. This makes the
failure survivable rather than preventing it.
2026-08-12 09:42:11 -04:00
jpolo1224 76af990eed VK: stop reserving a desktop-sized descriptor cache on Android
A heap profile of Ratchet & Clank on an Adreno 740 put the only real growth during
play on vk::descriptor_set: 56 sets created in half a session, 49MB, through
simple_array::reserve from descriptor_set::operator=. Nothing else grew that was
not one-time JIT or shader compilation.

Each set reserves m_pool_size entries in three pools the first time it is used --
16448 image infos, 16448 buffer infos, 16448 buffer views, about 920KB -- and there
is one set per shader program. simple_array::clear() only resets the size, so that
memory is held for the object's whole life, and the total climbs for as long as new
pipelines keep appearing. Ratchet compiles a lot of them.

It is also why this was invisible from the Vulkan side: these are plain malloc, not
device memory, so the VMM never sees them and no amount of texture eviction reclaims
them. The tester's log shows the shape exactly -- our pool at 516MB while the process
walked from 4448MB to 5626MB without ever dropping, then died on a 32MB allocation.

The reservation is a correctness requirement, not a tuning knob: push_*() hands
Vulkan the address of a pool entry and it has to stay valid until flush(), so the
pools must not reallocate while writes are pending. What makes it safe to shrink is
that max_cache_size is also the flush threshold, in both on_bind() and
storage_cache_pressure(), so the queue can never outrun the reservation. Moving it
takes the guard with it and leaves the same 64 entries of headroom.

1024 on Android: about 57KB a set instead of 920KB, for one extra
vkUpdateDescriptorSets per 1024 writes. Desktop keeps 16384.
2026-08-12 09:05:39 -04:00
jpolo1224 441c3f1bde Merge PR #34: vectorize the primitive-restart index upload on ARM64 2026-08-12 09:03:52 -04:00
Zulux91 fcdd7cd6af RSX: static_assert the all-ones identity the NEON restart lanes rely on
The vector body stores vorrq(v, eq) into restart lanes, which is all-ones
regardless of what index_limit() returns -- it only matches the scalar
tail's index_limit store because index_limit is all bits set. Assert that
beside the splats so a change to index_limit fails the ARM64 build instead
of silently diverging the vector body from its own tail. No codegen change
(emitted assembly is identical).
2026-08-12 07:49:32 -05:00
Zulux91 ad15f81d6e RSX: clarify comments on the ARM64 index-upload paths
Explain why the non-restart loop stays scalar (clang already
auto-vectorizes it) and spell out which allocations each caller passes,
so the no-overlap contract of upload_untouched_neon is checkable.
2026-08-12 07:24:20 -05:00
Zulux91 2df75aa604 RSX: vectorize the primitive-restart index upload on ARM64
The primitive-restart variant of upload_untouched had no SIMD path on
ARM64: the asmjit builder is x86-only, and clang cannot auto-vectorize
the scalar loop (-Rpass-analysis: "value that could not be identified
as reduction is used outside the loop") because the min/max updates are
conditional on the restart compare -- while the non-restart loop next to
it does auto-vectorize. Net effect: 16 scalar instructions per index on
a path some titles saturate. Measured on a Snapdragon 8 Elite (Odin 3),
Virtua Tennis 4 routes its entire indexed-draw traffic through this
loop: 2.81 billion indices in a 9-minute match session, median 159k
indices per frame.

Port the x86 lane algebra to NEON, 8x u16 / 4x u32 per iteration: the
restart-equal mask ORs the lane to all-ones for the min accumulator and
the store (all-ones is index_limit, exactly what the scalar loop
writes) and BICs it to zero for the max accumulator, so restart lanes
can never win either reduction; UMINV/UMAXV reduce once at the end and
the tail stays scalar. Baseline v8.0 AdvSIMD only.

Supporting results, all on the Odin 3 with the system driver:

- Correctness: 216-case differential (scalar vs NEON vs the dispatched
  path; every tail residue mod 8 and mod 4; restart index absent,
  present, 0, index_limit, all-restart, and index_limit present while
  not the restart value; u16 and u32) ran on device at RSX init in all
  four A/B runs: 0 mismatches. An independent 65,000-case host-side
  model of the same lane algebra also matched the scalar loop, and a
  blind review of the diff could not construct a diverging input.

- Performance A/B (cntvct_el0 around the dispatch, null-region
  calibration subtracted, per-window medians over 120-flip windows with
  >10k restart indices/flip, runs interleaved scalar/NEON/NEON/scalar):
    scalar: 0.02504 and 0.02464 ticks/index  (1.30 ns/index)
    NEON:   0.00346 and 0.00381 ticks/index  (0.19 ns/index)
  ~6.8x faster per index; scalar-scalar repeatability 1.6%. Worst
  single-frame cost in this loop fell from 3.64 ms to 0.99 ms. FPS
  stayed 60/60 in all runs on this device; the win is RSX-thread
  occupancy and worst-frame cost, and would be frame time where the
  RSX thread is the bottleneck.

Titles that never enable primitive restart are unaffected: they route
through the untouched path, which clang already vectorizes.
2026-08-12 05:53:44 -05:00
jpolo1224 b0c7a0260e VK: judge memory pressure against what the driver will give, not our own cap
vmm_determine_memory_load_severity is a set of thresholds on get_memory_usage,
which is usage/budget straight out of vmaGetHeapBudgets. VK_EXT_memory_budget was
never enabled, so VMA had no budget from the driver and used the heap size -- and
where pHeapSizeLimit is set, that limit, which is our own vram_allocation_limit.
An Adreno 740 log shows the consequence: 516MB against a 2048MB cap is 25%, below
even the 50% mark, so the allocator kept its fastest flags, severity stayed 'low',
and the 75/90/95 eviction ladder never fired. The first allocation the driver
refused was also the first sign of trouble, and that one is fatal.

Enabling the extension gives VMA the driver's own estimate. VMA takes the smaller
of it and pHeapSizeLimit, so the cap still caps -- it just stops being mistaken for
headroom that exists. Gated on support and logged when absent.

NOT a fix for the Ratchet & Clank crash this was found in, and it should not be
credited as one. That log leaks about 60MB per sample, 4448MB to 5626MB with no
drop, while our own pool sits at 516MB -- so nearly all of it is outside anything
VMA can see or evict, and a truthful budget only makes us give up our own memory
sooner. The tester reports 0.4 unaffected, which makes it a 0.5 regression still
to be found; see the Adreno per-vkCmdEndRenderPass allocation already recorded
against this codebase.

Also: licences can be removed. Installing one was one-way -- the row existed only
to prove the install had happened -- so a wrong or duplicate .rap could only be
cleared through a file manager, which on a scoped-storage device most people
cannot do at all. Confirmed before deleting, like uninstalling a title, because
content stops working without it.
2026-08-11 23:43:23 -04:00
jpolo1224 7745a3c92d Apply the README trim from PR #28
Reverted during the merge on the assumption it was a contributor's deletion; it
was a deliberate one, so take it as authored.
2026-08-11 22:01:43 -04:00
jpolo1224 b09815e595 Pad: carry analog button pressure through to the game
Every pressure-capable button was fully digital. _rpcsx_overlayPadData ended with
btn.m_value = m_pressed ? 255 : 0, and that value is what cellPad copies into the
press byte a game reads for an analog button, so no half-press could ever reach
one. Rpcs3Bridge.setPadButton threw the magnitude away before that, using `range`
only for stick directions and calling applyButton -- pressed or not -- for
everything else.

Two features were silently dead as a result. A physical L2/R2 went 0 to 100 like
a digital button, reported on Iron Man, whose level-two hover tutorial cannot be
passed without a half-press; the trigger axis was read and scaled correctly all
the way to the JNI boundary and discarded there, which is why remapping and
recalibrating changed nothing. The touch overlay's pressure modifier had the same
end: it computes a range through pressureRangeFor and hands it to the same call.

Pressure now travels as its own export rather than widening overlayPadData, whose
signature is frozen -- the core is dlopen()ed and updated independently of the JNI
glue, so a wider existing export would have older glue passing a garbage argument.
Glue or core predating _rpcsx_overlayPadPressure keeps the old digital behaviour.

0 means "nothing analog drives this button", which is a safe sentinel rather than
a lost level: an unpressed button already reports 0, so a pressed button at 0
cannot occur, and the zero-initialised array is exactly the previous behaviour.
Pushed only when it changes, so an all-digital pad adds no JNI call per event.

All twelve buttons the PS3 pad reports pressure for, not just the triggers, since
the offsets are contiguous and cellPad already routes each one. sendTrigger also
floors to at least 1: the lightest real squeeze truncated to 0, which is the input
layer's "full press" convention and would have delivered the opposite of a
half-press.
2026-08-11 21:53:09 -04:00
jpolo1224 3791865e2f Merge PR #28: make a Vulkan driver that stops answering diagnosable
Rebased by the author onto 0.5, so the occlusion-query bound we shipped stays as
it is and this only adds diagnostics on top of it: the fatal throw that ended the
session is gone, and with it the Web of Shadows regression that kept both PRs out
of 0.5. Also leaves the wait on shutdown, so a driver that never answers cannot
wedge the exit.

README.md is deliberately not taken from the PR -- it removed the whole status
and differences-from-upstream section.
2026-08-11 21:47:58 -04:00
jpolo1224 6831c87a0e Merge PR #23: fix PS3 patch state reporting and verify downloaded patches 2026-08-11 21:47:43 -04:00
Zulux91 b4378d8977 Say why a custom driver cannot load, before trying to load it
adrenotools swallows this failure. When its dlopen of the custom driver fails it
logs to logcat and hands back the system driver, so the load looks successful
from here, the reason never reaches the emulator log, and the user runs a driver
they did not choose while believing otherwise. The existing dlerror() report
cannot fire, because the pointer that comes back is not null.

That cost real time. Mr Purple T29 fails on an Android 15 device with "cannot
locate symbol pthread_getaffinity_np", falls back, and every log looked exactly
like a successful custom-driver session -- I recorded it as passing a driver
comparison it had never taken. The only hint was its reported driver version
matching the system driver's, which took three saved logs side by side to spot.

The requirement is stated in the file. DT_VERNEED lists the libc versions a
binary needs, and T29 needs LIBC_36, meaning API 36, on a device that provides
35. Reading that before the attempt turns "failed to load" into the reason, and
covers the whole class of community drivers built against a newer NDK than the
device runs -- likely the most common way these packages fail.

Metadata cannot answer this: T29's own meta.json declares minApi 30. That field
is author-declared and unverified, so only the binary is trustworthy.

Reported through the emulator log as well as logcat. The UI glue can only reach
logcat, which is not the file anyone attaches to an issue -- the reason would
exist and no report would ever contain it. It lands beside the driver identity
that it explains.

Advisory on purpose. The load is still attempted and nothing is rejected, so a
wrong answer here costs one log line and never a working driver. It stays quiet
unless it positively finds a LIBC_<n> requirement above the running API, and
declines to answer at all when section headers are absent.

Verified on device both ways: T29 reports needing LIBC_36 against API 35 in
RPCSX.log, and stevenmxz v33, which loads correctly, produces nothing.
2026-08-11 19:59:51 -05:00
Zulux91 8fed09d40e Say which way an abandoned occlusion query was failing
Upstream now bounds this wait itself: warn at one second, abandon at
three and use whatever the query holds. That replaces the fatal timeout
this commit previously carried, and it is the better answer -- the throw
could end a session over a driver that was merely slow, at worst wrong
culling for a frame was the actual cost. What remains here is the part
the bound does not cover:

- On abandonment, ask the driver once more directly with
  VK_QUERY_RESULT_WITH_AVAILABILITY_BIT and log which way it is
  stalling: VK_NOT_READY, or VK_SUCCESS with the availability word still
  clear. The two are indistinguishable through poke_query and need
  different conversations with whoever maintains the driver.

- Leave the loop when emulation is aborting. A driver that never answers
  must not also wedge the exit path, and the value is irrelevant once
  the session is going away.

Found chasing a Skate 3 freeze on an Adreno 830, where stevenmxz's gen8
driver builds accept occlusion queries and never complete them; the
system driver completes them in microseconds.
2026-08-11 19:59:51 -05:00
Zulux91 8772786f9a Say which Vulkan driver actually answered
The startup log named the GPU and a driver version, and on Android neither
identifies the driver. adrenotools' hook falls back to the system driver when
its dlopen of the custom one fails, and reports that only to logcat, so a
session that silently ran the system driver logged exactly the same thing as
one that ran the custom driver it was asked for.

That is not hypothetical. Chasing a Skate 3 freeze on an Adreno 830 I recorded
a custom driver as passing a test it never took: it had failed to load with
"cannot locate symbol pthread_getaffinity_np", fallen back, and the log still
said the custom driver was bound. The only tell was that its reported version
matched the system driver's exactly, which needed three saved logs side by side
to notice.

Logs the driver identity Vulkan already reports -- name, driverID, info and
conformance version, all of which were being fetched and thrown away -- and
falls back to saying the identity is name-derived when VK_KHR_driver_properties
is missing, which is common on the older Android devices this matters most on.

Where a custom driver was requested and Qualcomm's own driver answered, that is
a silent fallback, since adrenotools installs Mesa/Turnip builds. It now says
so, and points at the logcat line carrying the actual reason.

The loader's own message no longer claims more than it knows: the handle it
binds is the one it was handed, and whether the driver behind it is the
intended one is not something it can see.
2026-08-11 19:57:37 -05:00
jpolo1224 5405d71e2e Update README.md 2026-08-11 18:12:01 -04:00
jpolo1224 93c6df7e42 Settings: clear the core tuning left pinned while debugging
Raw core overrides re-push after the curated settings, so a stale one silently
beats the UI with nothing on screen to explain it: the settings screen read SPU
Block Size = Safe for hours while config.yml read Mega.

Mega is the one that mattered. It produces very large compilation units, and those
are what fail AArch64 register allocation with "Cannot scavenge register without
an emergency spill slot" -- which is what put SPU threads on the interpreter
fallback at all. With it cleared, no block fails to compile and the fallback never
engages. Every "cannot be compiled" chased in these sessions traces back to it.

Cleared in every scope, because a title can pin a key the global also pins: Arkham
City carried Accurate SPU Reservations true as a raw per-title override against
false globally, so clearing one scope did nothing and the two readings looked
contradictory.

Per-title Accurate SPU Reservations values go too, except Web of Shadows, which is
the title it was measured on. Off is off-spec -- it forces the SPURS scheduler to
HLE and bypasses the reservation lock -- and a title left that way desyncs until
its SPU threads execute whatever they land on, which is how Arkham City ended up
dying with "Unknown STOP code: 0x0".

Adds CoreSettingOverrides.forgetEverywhere for the all-scopes case.
2026-08-11 15:54:54 -04:00
jpolo1224 f056d6fc86 Video: keep the VRAM heap cap at 2048
3072 was set to get the God of War 3 demo past an allocation failure, but that
failure was measured before the uninterruptible reclaim fix landed, and the cap is
not coordinated with the texture cache, which budgets itself up to 2560MB on
Android. Raising one without the other let the total grow with it: Batman: Arkham
City reached 5596MB resident against a 6246MB peak on a 7.2GB device and stalled
after a while, with no allocation failure to point at.

2048 is the value that shipped before, and it is where the sum of the two sat when
that game worked. Budgeting the cap and the cache together is the actual fix and
is not attempted here.
2026-08-11 15:10:41 -04:00
jpolo1224 ae1caf915b Release 0.5: purge the profiling overrides, raise the VRAM heap cap
The RSX profiler was still recorded as a raw core override from the debugging
work, so config.yml read "RSX Profiler: true" while nothing in the UI said so --
the same divergence as the relaxed-ZCULL one, since overrides re-push at the tail
of applyTo. It writes a bucket report every 300 frames and keeps per-scope timers
on the RSX thread, which is not something to ship enabled. The first purge had
already marked itself done, so this takes a new key.

VRAM allocation limit is applied as VMA's pHeapSizeLimit, which makes it a hard
ceiling rather than an eviction threshold: once total allocations reach it VMA
returns OUT_OF_DEVICE_MEMORY however much the device has free. Lowering it does
not make the cache release earlier, it makes allocation fail earlier. The God of
War 3 demo was measured failing a routine 24MB request at 1024 while the process
held 1.6GB resident and 280MB in that pool, and failing at 2048 one screen later.
3072 leaves the caches room while keeping the bound that stops an unbounded quota
driving the process to 4.3GB and getting it killed.

The allocation failure now names the request size and the cap alongside it, since
"Out of video memory" alone cannot separate a full device from an artificial
ceiling, and those need opposite fixes.
2026-08-11 14:06:11 -04:00
jpolo1224 9128a75c0e VK: reclaim what can be reclaimed when the renderer is uninterruptible
Refusing outright skipped the allocator's own recovery. That path is "if
OUT_OF_DEVICE_MEMORY and vmm_handle_memory_pressure(...) succeeds, retry the
allocation", so returning false meant the retry never ran and the allocation died
having freed nothing: God of War 3 reached it with zero reclaim attempts and zero
recoveries logged.

Only the fatal branch needs the queue idle, which is what the flush inside it is
for. The rest is reachable while uninterruptible: the texture cache purges its
unreleased pool, and at severe it also drops unlocked sections. RPCS3 already runs
exactly that with no flush whenever pressure is non-fatal, so this is the existing
contract rather than a new risk. Severity is clamped below fatal so the
flush-dependent path stays unreachable.

Measured after: eviction runs and reports releasing resources, and the allocator
retries. God of War 3 still fails, but now for the honest reason -- the device is
out of memory, with 123MB free of 7.2GB and the emulator resident at 4.3GB -- and
not because nothing was ever given the chance to run.
2026-08-11 13:55:43 -04:00
jpolo1224 785c4d5627 VK: decline video memory pressure instead of aborting, and budget it lower
on_vram_exhausted asserted that the renderer was interruptible. Eviction really
cannot run in that state, since it would touch resources the driver may still be
reading, but that is a reason to refuse rather than to kill the thread -- and
refusing is already the supported answer: the OOM path in VKDraw treats false as
using placeholder textures, which it notes can cause graphics glitches but
should not crash otherwise.

God of War 3 hit it by skipping the intro screens, which pushes a burst of surface
and texture allocation through a point where the renderer is uninterruptible. The
RSX thread died there, audio kept playing, and it presented as a hang. With the
refusal in place the same run reports the real problem instead:
VK_ERROR_OUT_OF_DEVICE_MEMORY from the allocator.

Which it genuinely is. VRAM allocation limit was also lowered from 2048 to 1024:
the first value was still above what the device could give us -- 5355MB resident,
99MB free of 7.2GB -- so the budget was never reached before the system ran dry,
which defeats its only purpose. It has to sit below what allocation can actually
satisfy, so eviction starts while there is still room to allocate.
2026-08-11 13:45:54 -04:00
jpolo1224 38424a59bd SPU: mark failed program ranges; Video: budget VRAM on mobile
Three things, all found by measurement after the interpreter fallback started
being used in anger.

Marking only the entry point made the interpreter release the thread after one
instruction, whereupon the recompiler tried the next address, failed the same way
and marked that too. 111 consecutive entries were recorded walking two blocks four
bytes at a time, each step paying a full failed LLVM compile. The failed set now
holds ranges, so a thread stays interpreted for the whole block and leaves when
execution genuinely moves past it: 111 markings became 1.

The range test then ran per interpreted instruction and took a reader lock each
time, which put shared_mutex::imp_lock_shared at 28% of the whole process against
23% for the interpreter itself. The extent is now cached on the thread when the
fallback engages, so the loop compares two integers.

The switch was also logged once per thread, but the flag is cleared on every exit,
so the guard fired on every re-entry: God of War 3 wrote thousands of lines a
second ping-ponging between two addresses. Removed; the block is still recorded
once when it is marked.

Separately, VRAM allocation limit was left at upstream's 65536 MB, which means no
limit and assumes a discrete card. Here the GPU shares system memory with the OS
and our own host allocations, so the texture cache is never asked to evict and
grows until allocation fails -- and failing is fatal: God of War 3 dies in
on_vram_exhausted on ensure(!vk::is_uninterruptible() && ...), because VRAM ran
out where the renderer cannot safely evict. Measured at the crash: 5355MB
resident, 99MB free of 7.2GB. 2048 leaves room for the guest's own memory, the
host caches and the OS.
2026-08-11 13:39:05 -04:00
jpolo1224 8c707648cb SPU: leave the interpreter once the uncompilable block is behind us
The fallback flag was set once and never cleared, so a thread that met a single
block it could not compile interpreted everything it ran from then on. Correct,
but these are SPURS kernels doing real work, and Sonic Unleashed reached its
loading screen that way and then crawled through it.

The failed set holds entry points, so this keeps interpreting while pc sits on the
bad entry -- which is where a branch-to-self idle loop stays -- and releases the
thread as soon as execution moves past it. Only the block that cannot be compiled
is interpreted; the rest of the thread runs recompiled.

Leaving is safe at any instruction boundary, since all SPU state lives in
spu_thread, which is the assumption the JIT dispatch already makes. Re-entering
the bad block sets the flag again.
2026-08-11 12:55:02 -04:00
jpolo1224 9543378661 SPU: route the uncompilable-block fallback to the interpreter that works
A block that fails to compile switches its thread to the interpreter. On ARM64
that fallback called spu_runtime::g_interpreter, which with a recompiler selected
is the LLVM-built interpreter, and calling it there executes nothing: measured a
million consecutive calls on Sonic Unleashed's stuck SPURS kernel without pc
moving once. The thread then spins in that loop forever at a fixed pc with no
flags set, which reads as a busy SPU and hangs the title with no diagnostic at
all. Any block that fails to compile landed there, so this was not one game.

old_interpreter is what the static decoder ultimately runs, through
tr_interpreter, and it is self-contained -- opcode table, thread, local store. Its
static-decoder-only check rejected exactly the case that needs it, so it now also
accepts a thread already marked for fallback.

Getting there also needed the give-up paths fixed: the TBL2/TBX2 retry could
return null with an empty error and fall through every branch unmarked and
unlogged, so nothing recorded that a block had been abandoned.

The stall dump now carries SPU event, MFC and interrupt state, which is what made
this findable: parked kernels showed pending=0 (no lost wakeup), intr_en was 0 on
healthy threads too (not interrupts), mfc_q was 0 everywhere (no stuck transfer),
and interp_fb=1 on the frozen thread pointed at the fallback itself.
2026-08-11 12:44:04 -04:00
jpolo1224 0b5a43c602 vm: report who a stuck writer_lock is waiting on
Both waits in writer_lock are unbounded and silent. The acquire loop spins until
every range lock bit clears, and the range_lock path then spins until every
registered PPU thread reaches cpu_flag::wait. A thread that never gets there hangs
every other thread that takes a reservation, and leaves nothing behind: from
outside it reads as a clean guest deadlock with everything in a legitimate wait.

Both now log once, far past any plausible contention, naming the held range locks
or the PPU thread being waited on.

They paid for themselves immediately on Sonic Unleashed, which deadlocks at the
SEGA logo. Both stayed silent across several boots, which ruled out the VM lock
entirely -- worth having, since main_thread was pinned in cellSpursRemoveWorkload
carrying cpu_flag::memory without cpu_flag::wait, which looks exactly like this
bug and is not. The game hangs in a different state on different boots, so it is a
race elsewhere in SPURS.
2026-08-11 10:50:08 -04:00
jpolo1224 5997681fed Packages: install from storage that cannot be read directly, and name what is installed
Issue #16, both halves.

Installing a .pkg or .rap off a USB-OTG drive already went through the system
picker, and the descriptor it returns is handed to the native installer as a raw
fd, so a 40 GB package costs no copy. That holds only while the provider is
backed by real storage. The third-party USB-OTG and cloud apps people reach for
when the platform will not mount their drive return a PIPE, and every install
entry point seeks -- getFileType sniffs the magic and rewinds, package_reader
jumps around the archive -- so lseek failed with ESPIPE and a perfectly good
package was reported as unsupported or broken. Those descriptors are now
detected with the same lseek the core will make, and only those are copied to
real storage first, onto whichever of the emulator's own storage and the app
cache has more room. The copy is checked against the size the provider reported:
a short copy does not throw, it produces a truncated package that fails much
later as "broken", which reads as a bug report about the package.

Split releases picked through the system picker arrived in the order the user
tapped them, and installSplitPkg takes the order given as the part order, so
picking part 2 first extracted into a broken install rather than failing. Both
pick paths now sort the parts, digit runs numerically, since plain string order
puts part 10 between part 1 and part 2.

Installed titles were listed by title id alone -- NPUB90434, BLES01807 -- next
to an Uninstall button, which is where it hurt most: choosing which of two demos
to reclaim space from meant looking the ids up elsewhere. TITLE now comes out of
the install's own PARAM.SFO, read off disk rather than through the library cache
so a title the scanner has not seen yet is still named. The id stays on a second
line because patches, cheats and compatibility lists are keyed by it. Licence
files carry the same id inside their content id, so a .rap can name the game it
unlocks instead of being one of a row of indistinguishable hex strings.

The SFO field reader is the scanner's CATEGORY reader generalised rather than a
second copy of the 16-byte index-entry layout.
2026-08-11 10:43:11 -04:00
jpolo1224 49012ba969 Settings: seed per-title core settings, starting with Web of Shadows
Accurate SPU Reservations off is worth a large amount in Spider-Man: Web of
Shadows and is not safe globally, so it goes in that title's own override rather
than the default.

Its SPURS reservation traffic serialises behind the global exclusive
vm::writer_lock that every reservation_op takes, which no amount of CPU can help:
all six SPU threads and several PPUs were measured yielding at the same rate with
18.8% of total CPU in sched_yield. Off, SPURS takes the lock-free path and
vm::writer_lock fell from 8.06% to 0.96%.

Kept per-title because it is off-spec -- upstream defaults it on, and Sonic
Unleashed fails EARLIER with it off, reaching neither the loading icon nor the
logo, which is consistent with the bypass being the SPURS area itself.

The seed writes only fields a title does not already carry, so a deliberate
change is never overwritten, and it runs once. Other regions of the same game
need their own entry.
2026-08-11 10:26:35 -04:00
jpolo1224 0909f77a95 RSX: charge the empty-ring yield to idle, and stop draining the present queue on flush
Two things, both about the RSX waiting rather than working.

flush_command_queue ended by draining the present queue in case a queued frame
still held a ref to the command buffer just taken. It cannot: next() hands them
out from a 512 entry ring and the queued list is bounded at flip to
m_max_async_frames - 1, so the buffer being reused is hundreds of frames retired.
The guard was unreachable and the cost was not -- check_present_status pokes the
oldest queued frame's swap command buffer, and on Adreno vkGetFenceStatus blocks
until signalled rather than returning VK_NOT_READY, so a poll written to be cheap
became a full GPU sync. 1.32 times a frame at about 11ms: Fence poll 14.6ms ->
0.033ms, frame 44.5ms -> 36.6ms. Same fault as the two sites removed earlier;
this one sat inside flush_command_queue rather than on the present path. Ruled
out first: identical frame time at quarter resolution, and forcing the swapchain
pre-transform to match the surface left it unchanged.

The empty-ring yield is now charged to idle. It sits inside fifo_decode, which is
the enclosing scope of the whole run loop, so waiting on an empty ring was
reported as decode work -- Idle 0.003ms against FIFO decode 20.4ms, while a
native profile of the same thread put 34% of its cycles in sched_yield. The
bucket report and the profiler disagreed and the bucket report was wrong, which
has now produced two wrong conclusions in one session.
2026-08-11 10:13:53 -04:00
jpolo1224 0cf85cb902 RSX: drop the FIFO idle sleep; log the swapchain pre-transform
The 50us backoff in the FIFO_EMPTY path was added when the RSX thread was
measured spending 66% of its cycles in sched_yield on a machine starved for
cores: the affinity mask confined six SPU threads to four cores, and the
reservation path serialised everything behind a global lock, so a spinning RSX
took a core from threads that needed it.

Neither holds now, and the trade inverted with them. Measured after both were
fixed: 34% of eight cores busy, two to four threads runnable, five idle. Nothing
wants the core the sleep gives back, and the RSX sits on the frame's dependency
chain, so sleeping only delays it noticing the guest has produced work.

Also log the surface transform at swapchain creation. 30% of the frame is now in
check_present_status waiting on acquire_next_swapchain_image, which is not GPU
work -- a quarter-resolution run measured the same frame rate. Declaring IDENTITY
while the surface is rotated hands the rotation to the compositor, which can hold
images longer before releasing them for acquire. Logged rather than changed:
matching currentTransform means applying the rotation ourselves across the blit
and the overlay pass, and that is only worth doing if the two actually differ.
2026-08-11 10:02:54 -04:00
jpolo1224 cb3670e2db Settings: restore upstream's GETLLAR busy-wait percentage
Dropped to 20 while the emulator was starved for cores, reasoning that a spinning
SPU steals a core from threads doing real work. Two things have changed
underneath that: the affinity mask no longer confines six SPU threads to four
cores, and the reservation path no longer serialises everything behind a global
lock. Measured after both, in game: 34% of eight cores busy, two to four threads
runnable, five idle, and no thread near saturation.

Tested at 100 and at 20 with no difference, which fits -- the wait is no longer on
the critical path, so how it waits does not matter. Upstream's value stands
rather than carrying a divergence that buys nothing.
2026-08-11 09:52:02 -04:00
jpolo1224 5f346b5e2c Settings: default the thread scheduler back to the OS
The affinity migration turned the scheduler on so the big.LITTLE mask would
apply, keeping SPU and RSX off the A510s that run at roughly 27% of prime-core
capacity. That reasoning holds for one thread per core and breaks down at six.

Measured in game on a Snapdragon 8 Gen 2, reading the masks the threads actually
carry:

    app cpuset (top-app)  0-7    Android grants every core
    SPU[0..5]             3-6    six threads, four cores
    rsx::thread           3-7

Six SPU threads sharing four cores get about two thirds of a core each, which is
worse than one thread owning an A510 outright, and it caps the emulator: the
device sat at 60% busy with cores 0-2 idle while frames were slow. Spider-Man:
Web of Shadows is visibly better at OS.

Worth being clear that the mask is ours and not Android's -- the app is in
top-app with all eight cores granted -- which is also why these devices are
reported to run better under native Linux, where no such policy is applied.

The other modes stay selectable for anyone whose device disagrees.
2026-08-11 09:33:32 -04:00
jpolo1224 1c2d44502c cellPad: keep pressure values in the buffer when press mode is off
A DualShock 3 sends pressure bytes in every packet. The press setting governs how
much of the buffer the game is told is valid, not whether the pad produced the
values, so clearing the area diverges from the hardware: a game that reads a
pressure byte without having asked for press mode gets 0 where it would see a
press on a console.

Spider-Man: Web of Shadows does exactly that for R2. Tracing both entry points
showed it never calls cellPadInfoPressMode or cellPadSetPressMode, so the setting
stays at 0, yet it reads the R2 pressure byte to decide whether the trigger is
held. R2 did nothing in that game while every other button worked, from the
controller and from the touch overlay and after remapping to a different physical
button, because the digital bit was delivered correctly the whole time and was
never what the game looked at.

len is unchanged, so a game that honours it sees what it saw before.
2026-08-11 09:22:58 -04:00
jpolo1224 ba5e4ebd66 SPU: stop an out-buffer verdict from spinning GETLLAR forever
The out-buffer check answers 'unlikely to be a loop', and that answer is not
free. It resets the spin count and leaves the busy-waiting switch at umax, so the
caller skips busy_wait, skips the sleep path, and returns immediately: the SPU
re-executes GETLLAR at full rate with no backoff. The spin count never reaches 4,
so the spin optimisation is never evaluated, and the 400ms fallback that would
force a sleep is never reached either. One SPU in that state holds a core flat
out, and the setting meant to control this has nothing to act on.

Spider-Man: Web of Shadows sits in that case. Its GETLLAR sites use an LSA in the
top 64K of local store, which is what the check looks for, and process_mfc_cmd
measured 55% of all CPU across the process while the game ran at 10-15fps.

Re-entering the same site with the same stack 32 times is itself the evidence
that it is a loop, whatever the LSA looks like. After that the verdict is dropped
and the normal spin detection decides between busy-waiting and sleeping. Any real
change of site or stack resets the count, so a genuine OUT buffer still gets the
original treatment.
2026-08-11 00:01:45 -04:00
jpolo1224 a873ae8250 Settings: let idle SPUs sleep instead of spinning on a reservation
SPU GETLLAR Busy Waiting Percentage defaults to 100 upstream, meaning always
busy-wait. That suits a desktop, where the SPU threads have cores of their own
and spinning costs nothing else. Here six of them share eight cores with the PPUs
and the RSX, so a spinning SPU takes a core from the threads doing the work.

Measured on Spider-Man: Web of Shadows: process_mfc_cmd accounted for 55% of all
CPU across the process, and making its inner loop cheaper did not move the frame
rate -- the loop just ran more iterations in the same wall clock. That is what
identified it as a spin rather than as work, after two rounds of optimising the
iteration itself.

20 still favours a short busy-wait, so a reservation that frees quickly is caught
without a scheduler round trip, and only a wait that history says is long goes to
sleep. A deliberate change in All Core Settings still wins, since core overrides
replay after this.
2026-08-10 23:56:15 -04:00
jpolo1224 e8c499056b SPU: memoise the GETLLAR out-buffer check on its inputs
Gating the check on getllar_spin_count was not enough. That counter is reset from
several other paths, so it is frequently zero and the callstack was still rebuilt
constantly: measured 19.6% of all CPU inclusive in dump_callstack_list, the
largest single item after process_mfc_cmd itself.

Key on the values the answer actually depends on instead -- pc, the stack pointer
and the link register -- and recompute only when one of them moves. Only the
innermost frame is ever used, so that is all the memo keeps.

A stale answer across an unrelated LS write is acceptable here. This decides only
whether the address looks like a caller's OUT buffer, on a heuristic whose own
comment calls it 'unlikely to be a loop'.
2026-08-10 23:52:23 -04:00
jpolo1224 88f4162883 SPU: evaluate the GETLLAR stack heuristic once per spin sequence
The out-buffer check in the GETLLAR spin detector rebuilt the callstack on every
iteration of a busy-wait loop. dump_callstack_list walks the stack and calls
is_exec_code for each candidate, which allocates a vector<bool> and scans for
branch targets, so the cost is large next to what it decides.

On a whole-process profile of Spider-Man: Web of Shadows those three came to
about 14% of all CPU -- more than the RSX thread spent on the frame -- because
the game's SPU code spins on GETLLAR with an LSA in the top 64K of local store,
which is exactly the case the check looks at.

Once per sequence is enough. pc, ch_mfc_cmd.lsa, gpr[1] and addr are all compared
against the previous iteration a few lines above and any change resets the
sequence, so the callstack cannot move underneath a spin.
2026-08-10 23:43:41 -04:00
jpolo1224 771a0dc74e RSX: back off to a sleep when the FIFO ring has gone quiet
sched_yield does not idle a core. With every core already busy it returns almost
immediately and the RSX thread takes it again, running flat out producing
nothing. A native profile of Web of Shadows put 66% of this thread's cycles in
sched_yield and its kernel path against 9% in run_FIFO, which the bucket profiler
reports as a busy RSX because the yield happens inside the fifo_decode scope.

That is not free even with an empty ring. The RSX affinity mask covers the whole
fast cluster while the SPU mask is that cluster minus the prime core, so any of
this that lands off the prime core is taken from the SPU threads, and those are
what the frame is actually waiting on at 61% of all CPU.

The spin still runs 64 times before sleeping, so a producer that is merely slow
is met without a scheduler round trip. Only a ring that has genuinely gone quiet
reaches the 50us sleep, which is far below the frame times where this matters.
Android only.
2026-08-10 23:37:48 -04:00
jpolo1224 1458d5f2d9 VK: end open occlusion queries in end_renderpass, for every caller
A query that begins inside a render pass instance has to end inside that same
instance. Ending the pass underneath an open one leaves it permanently
unavailable, and on Turnip it takes the device with it, reported later against
poke_query because that is the first call that waits on a result.

Queries do begin inside render passes here. VKDraw only lifts them out when
use_strict_query_scopes() is set, and that is wired to Strict Rendering Mode, a
user performance setting rather than a driver quirk, so it is off for almost
everyone.

Twenty-one call sites end a render pass and only one, in VKDraw, ever paired
itself with a cleanup. change_image_layout alone ends 41 passes a frame in Web of
Shadows, and any of them can land while a query is open, which is why fixing the
two sites in the query pool moved the device loss from one minute to nearly three
rather than removing it. Holding the invariant in end_renderpass covers all of
them, including any added later.

The VKDraw site now cleans up before the pass ends rather than after, which is
the order the spec asks for; its own call becomes a no-op.
2026-08-10 23:27:17 -04:00
jpolo1224 2cd341e257 Settings: purge the stale raw Relaxed ZCULL Sync override
The migration that turned relaxed ZCULL on recorded it twice: once in the curated
store and once as a raw core override. The migration that turned it back off only
corrected the curated field, so the two stores disagreed, and the override is the
one that reaches the core -- overrides re-push at the tail of applyTo, after the
curated store has written the setting.

The toggle therefore read OFF while config.yml read 'Relaxed ZCULL Sync: true' on
every boot, with no way to change it from the UI. That is not cosmetic: relaxed
sync is what allows queries to be read while still pending, which is the path
behind the 'Dubious query data pushed to cond render' warnings, and it also
selects emulated predication in the VK backend.
2026-08-10 23:17:53 -04:00
jpolo1224 fa97d01b0d VK: do not emulate predication where we disabled the extension ourselves
Emulated conditional rendering exists for hardware that never had the extension.
Turning the extension off as a driver workaround enabled it by accident, because
both are selected by the same test, and the two halves do not fit together:
begin_conditional_rendering returns early without building m_cond_render_buffer,
while the vertex shader still reads that buffer at offset 0. It gets a zeroed
scratch buffer, predicates every draw away, and the game renders black with audio
and overlays still running. The all-ones word that disables predication sits at
offset 4 and is never reached, since the fallback leaves hw_cond_active set.

Off on these drivers means occlusion results stop culling draws, which is the
trade the workaround already documents.

The vendor comes off the GPU rather than from get_driver_vendor(), whose cached
value is not assigned until later in the same function and would still hold the
previous device's.
2026-08-10 23:14:07 -04:00
jpolo1224 703d98ae9f VK: disable conditional rendering on Turnip as well as Adreno
The existing gate covered only the proprietary driver and said Turnip was left
alone until there was evidence about it. There is now.

Web of Shadows loses the Vulkan device about a minute into gameplay on Turnip 26
/ Adreno 740. The assertion names poke_query, which is the first call that reads
a result rather than the one at fault. Conditional rendering is the only place we
record vkCmdCopyQueryPoolResults with VK_QUERY_RESULT_WAIT_BIT, and that form
makes the GPU block until the query resolves, so a query that never resolves
hangs the device instead of the caller and the watchdog ends the session. The
same run logged 169 'Dubious query data pushed to cond render' warnings, which is
this code being handed queries that are still pending.

It is also the churn: the aggregation barriers closed 42 of the 91 render passes
in a measured frame, and ending a pass on a tiler costs a tile store and reload.

Both drivers now fall back to thread::begin_conditional_rendering, the path
desktop already takes wherever the extension is absent.
2026-08-10 23:10:32 -04:00
jpolo1224 cf2481fc11 VK: close open occlusion queries before ending the render pass
Fixing the device loss by ending the render pass before vkCmdCopyQueryPoolResults
introduced a hang in its place. A query that begins inside a render pass instance
has to end inside that same instance; ending the pass underneath an open one
leaves it permanently unavailable, so get_query_result spins on poke_query with
no way out and the RSX thread stops.

Nothing reported it. The submit-time ensure() only checks that the query was
closed, and end_occlusion_query closes it a moment later, so the assert passes
while the result never arrives. The stall detector runs from do_local_task in the
FIFO loop, which the spin has already left, so the profiler charged the wait to
FIFO decode and the frame read as CPU-bound -- 93% in a bucket that was really
the thread sitting in sched_yield. Web of Shadows locked up this way after
reaching gameplay, audio and vblank still running.

Both sites that end a pass from the query path now close an open query first,
which keeps begin and end within one pass. The query is cut short, as it is
anywhere do_query_cleanup is used.

The wait itself is now bounded as well. It warns at one second and abandons at
three, using whatever the query holds: wrong culling for a frame is a better
failure than a thread that never returns, and the log names the cause.
2026-08-10 23:01:36 -04:00
jpolo1224 8041edf5bc RSX: publish GET only when it has advanced
The drain fix made every path that can idle or block publish GET immediately,
which is required for correctness: a producer waiting on ring space needs to see
the progress we made before we stopped consuming.

It publishes far more often than that requires. The empty and busy cases return
straight to the run loop, so a ring that has gone quiet re-enters them once per
iteration with GET unmoved. Web of Shadows measured 137000 loop iterations per
frame against 46000 method dispatches; the remaining 91000 were republishing a
value the guest already had.

GET shares a 64-byte line with put, which the guest PPU writes from another
cluster, so each of those is a coherence miss taken against the thread feeding
the ring. The cost lands on the producer rather than on the RSX, which is why it
presented as a freeze with sound still playing: the PPU stalls on the contended
line while threads that never touch it keep running.

GET is ours to write, so tracking the last published value and skipping an
unchanged store keeps the guarantee -- progress is still announced exactly once
after the last advance -- without the repeats.
2026-08-10 22:50:14 -04:00
jpolo1224 d909537f3f Stop copying query results from inside a render pass
vkCmdCopyQueryPoolResults must be recorded outside a render pass instance. This
recorded it inside one, with a comment saying we are technically supposed to stop
the pass first but that it does not matter on IMR hardware. It is not a
technicality -- inside a pass it is undefined behaviour, and a desktop GPU
tolerating it says nothing about a tiler.

This device lost the Vulkan device over it. The fault surfaced later, in
poke_query, because that is the first call that waits on a GPU result, so it read
as the query READ being at fault when the damage was done at record time. Only
the RSX thread died, so the process kept running with audio and vblank alive and
it presented as a hard freeze rather than a crash. Verified gone: zero device
losses on a run that previously died within a minute.

The pass is ended only when one is actually open, on a path that already stalls
for a GPU result, so the flush the upstream comment worried about is paid where
we were blocking anyway -- and disabling occlusion queries is not the
alternative, measured here at 80ms frames with broken visuals.
2026-08-10 22:34:20 -04:00
jpolo1224 0cbab0c3ae Join the GPU and CPU pass tables on the same ordinal, and reach external storage
The two by-pass tables were joined on counters that reset at different points.
tick_frame runs from on_frame_end, before flip; the GPU timer rotates its slot at
the top of flip and then drops every non-frame region on the fresh slot, which is
flip's own overlay and calibration passes -- and those still incremented the CPU
counter. So the CPU ordinal ran ahead by the number of present-path passes and
the two tables described different passes. A whole anomaly came out of that: a
pass whose GPU cost was joined to a neighbour's workload read as 36x the per-draw
cost of its peers. The comment claiming both reset on the same boundary was
wrong. Reset where the GPU slot actually rotates instead.

Also adds a Storage Access Framework route to the package installer. The in-app
browser walks java.io.File, which only reaches storage this process can open by
path, so a .pkg on a USB-OTG drive or an SD card was unreachable and had to be
copied to internal storage first. Packages are handed over as the descriptor SAF
already returned -- the native side takes a raw fd, so nothing is copied and a
4 GB package costs no extra space; licences are 16 bytes and their installer
wants a real file, so those alone are staged.
2026-08-10 22:28:33 -04:00
jpolo1224 3338eedfcd Bound the PPU compile serialisation so a dead worker cannot strand it
The low-memory serialisation added three days ago holds a std::mutex across the
LLVM compile itself. A worker that hits LLVM's fatal handler leaves through
pthread_exit, and bionic unwinds nothing on that path, so the mutex stays locked
by a thread that no longer exists and every remaining worker waits on it for the
rest of the session. Memory only falls as modules accumulate, so the tight-memory
branch is likeliest late in a run -- which is why it reads as the PPU cache
getting stuck at the very end, and why dropping to the interpreter avoids it.

Reported as Saint Seiya: The Sanctuary never finishing its module cache. A claim
taken by compare-and-swap and waited on with a timeout costs a stranded claim a
wait rather than the session; the memory back-pressure either side of it is
unchanged. Third time this fork has been bitten by an unbounded wait around a
thread bionic can kill without unwinding.
2026-08-10 22:13:01 -04:00
jpolo1224 062cde277a Stop the save-data list aborting where there is no media backend
overlay_audio.cpp already accounts for a platform with no video source; the same
ensure() was left in overlay_video.cpp. Android's make_video_source returns
nullptr, and overlay_save_dialog builds a video_view for EVERY entry on all three
of its paths, so opening a save list aborted as soon as there was one save to
draw. It presents as the save menu never opening -- reported against Ratchet &
Clank: Tools of Destruction and Devil May Cry 4, and against Web of Shadows,
which stalls only once a save exists to be listed. Bundling the overlay icons
was necessary but not sufficient: the dialog still could not survive drawing.

The still image is what an entry needs; the animated ICON1.PAM is the part no
backend here can supply. Also dumps SPU thread pc and block hash alongside the
PPU dump when frames stop, which is what named the SPURS kernels as idle rather
than spinning in guest code.
2026-08-10 22:09:59 -04:00
jpolo1224 7b0f1cd6de Ship the overlay icons the native UI has been drawing without
overlay_controls.cpp loads a fixed set of PNGs -- button glyphs, save.png,
new.png, spinner -- through fs::get_config_dir() + Icons/ui/. Desktop ships them
beside the binary; nothing put them on Android, so every load failed and the log
said so on each boot. The visible cost was cellSaveData's list: it is a native
overlay that draws its rows with save.png/new.png, so the load-save menu a game
opens never appeared. Reported against Ratchet and Clank: Tools of Destruction
and Devil May Cry 4, both fine on emulators that ship the icons.

Bundled from bin/Icons/ui and staged into config/Icons/ui once, revision-guarded,
before the core can draw its first overlay.
2026-08-10 18:36:17 -04:00
jpolo1224 05842b3115 Retranslate every language from the current English map
The nineteen translation files were ARMSX2-era: about nine hundred of their keys
still existed and showed the old PS2 wording -- worse than the English fallback,
which is at least right -- and roughly eight hundred current keys had no
translation at all. Regenerated all nineteen from the 1041-key map, batch plus
per-line retry, with every %s/%d checked against the source so no broken format
string ships. A string that would not translate is omitted and falls back to
English rather than shipping wrong.
2026-08-10 18:36:17 -04:00
jpolo1224 d6235c5802 Remove the PS2 leftovers a PS3 emulator was still carrying
The memory-card and PNACH patch screens were PS2 concepts with no PS3 counterpart
and no caller left -- the drawer had already been cleaned, so they were dead code
holding dead strings. The PNACH downloader went with its only consumer, and the
PCSX2-Android.ini seed could never exist under this package. The session log now
announces ARMSX3_INIT instead of PCSX2_INIT, which had every bug report opening
with the name of a different emulator.

The English map drops 46 PS2 strings and 619 orphans nothing references (1705 ->
1041 keys), rewords the four live strings that still said memory card, and renames
about.pcsx2.* to about.rpcs3.* to match what they already said.
2026-08-10 18:36:17 -04:00
Zulux91 4101367b2d Verify downloaded patches and match the engine's wildcard serial
Three things in the patch download path, all of them things desktop
RPCS3 already does.

The download URL had the patch schema version written into it as 1.2.
That is right today, but patch_engine::load rejects any file whose
Version header doesn't match the core's patch_engine_version, so the day
upstream bumps that constant every download starts failing to parse. The
core now hands the version out through patchEngineVersion() and the URL
is built from it. I also check the version the server echoes back, which
turns a several megabyte download into an early error instead of a
parser complaint.

Nothing verified the sha256 the server sends alongside the patch text.
Desktop checks it before it writes anything (patch_manager_dialog::
handle_json). Patches are writes into the guest executable, and
move_file/hide_file patches reach the emulator's own filesystem, so I'd
rather not import bytes that aren't what the server hashed. Mismatches
get their own message rather than being reported as a parse failure.

Last, the wildcard serial. patch_key::all is spelled "All", and
patchSetEnabled compared against a lowercase "all", so a patch carrying
a wildcard entry never had that entry written.

I first "fixed" the same typo in patchesList and let per-game lists match
the wildcard too. Ooops. Turns out that is not a typo doing nothing, it
is a typo doing the right thing by accident: wildcard patches are keyed
by SPU or PPU hash and leave the serial as "All" because the hash is the
filter, so they belong to no single game. There are 17 in the database,
and on device Skate 3 cheerfully offered me a pile of LittleBigPlanet
MLAA patches, where toggling one writes the shared entry and changes
every other game as well. Desktop shows them once under an "All titles"
node, so per-game lists stay serial-only here and the global list is the
equivalent. Only patchSetEnabled and the enabled-state read get the
spelling fix.
2026-08-10 05:48:03 -05:00
Zulux91 466d85d5b9 Report PS3 patch state per game and say when it takes effect
Two things in the patch list reported something the game wasn't getting.

First, patchesList marked a patch as enabled if any entry under its hash
was enabled, whichever serial that entry belonged to. patchSetEnabled
writes per serial, so a patch I switched on from one game's list showed
as on in every other game that patch covers. It now scans only the
requested serial, plus RPCS3's "all" wildcard, which really does apply to
the game being listed.

Second, the patch engine builds its map once while the game loads, then
writes the patches into each module as that module loads. Toggling a
patch only rewrites patch_config.yml, so nothing happens in a game that
is already running. The in-game tab never said so, which makes the switch
look broken. It says so now, above the list.
2026-08-10 05:01:54 -05:00
207 changed files with 29320 additions and 19061 deletions
+3
View File
@@ -112,3 +112,6 @@
path = 3rdparty/protobuf/protobuf
url = ../../protocolbuffers/protobuf.git
ignore = dirty
[submodule "3rdparty/oboe/oboe"]
path = 3rdparty/oboe/oboe
url = https://github.com/google/oboe.git
+15
View File
@@ -141,6 +141,12 @@ else()
add_subdirectory(cubeb EXCLUDE_FROM_ALL)
endif()
# Oboe (Android only)
if(ANDROID)
message(STATUS "Using static oboe from 3rdparty")
add_subdirectory(oboe EXCLUDE_FROM_ALL)
endif()
# SoundTouch
add_subdirectory(SoundTouch EXCLUDE_FROM_ALL)
@@ -367,6 +373,15 @@ add_subdirectory(fusion EXCLUDE_FROM_ALL)
# FERAL INTERACTIVE
add_subdirectory(feralinteractive EXCLUDE_FROM_ALL)
# LSFG: Lossless Scaling frame generation. Android only, and deliberately NOT EXCLUDE_FROM_ALL --
# libarmsx3_lsfg.so has to be built and packaged even though nothing links it, because the core
# reaches it by dlopen rather than by linking. Marking it excluded produces a build that succeeds
# and an APK with no frame generation in it.
#
# The subdir returns immediately when the submodule is absent, so a checkout without it still
# builds; frame generation simply reports itself unavailable at runtime.
add_subdirectory(lsfg)
# add nice ALIAS targets for ease of use
if(USE_SYSTEM_LIBUSB)
add_library(3rdparty::libusb ALIAS usb-1.0-shared)
+46
View File
@@ -0,0 +1,46 @@
diff --git a/llvm/lib/Target/AArch64/AArch64FrameLowering.cpp b/llvm/lib/Target/AArch64/AArch64FrameLowering.cpp
index d89a972f5d..f64e551a51 100644
--- a/llvm/lib/Target/AArch64/AArch64FrameLowering.cpp
+++ b/llvm/lib/Target/AArch64/AArch64FrameLowering.cpp
@@ -2504,8 +2504,40 @@ void AArch64FrameLowering::determineCalleeSaves(MachineFunction &MF,
RegScavenger *RS) const {
// All calls are tail calls in GHC calling conv, and functions have no
// prologue/epilogue.
- if (MF.getFunction().getCallingConv() == CallingConv::GHC)
+ if (MF.getFunction().getCallingConv() == CallingConv::GHC) {
+ // ...but they can still need an emergency spill slot.
+ //
+ // Returning here skips every path below that reserves one, so a GHC function never gets
+ // a scavenging frame index on AArch64. That is safe only while the premise holds. It
+ // stops holding as soon as the allocator spills: the function then has real stack
+ // objects, eliminateFrameIndex may need a scratch register to materialise an offset,
+ // and GHC has reserved nearly every GPR, so there is no free register to take and no
+ // slot to spill one into. The scavenger then aborts the whole module with
+ // "Cannot scavenge register without an emergency spill slot".
+ //
+ // Reproduced with RPCS3's PPU recompiler, which emits ghccc for every guest function.
+ // A single function of Saint Seiya: The Sanctuary (BLES01421) fails this way, and losing
+ // it costs the entire module, whose functions then fall back to an interpreter loop. The
+ // failure needs ghccc AND a scheduling model that pushes pressure over the line (it
+ // reproduces on cortex-x1/x2/x3 and cortex-a55, not on cortex-a76/a78/generic) AND -O2;
+ // remove any one and the same function compiles.
+ //
+ // Gated on the function actually having a frame, so a GHC function with no stack objects
+ // still gets no prologue and nothing changes for it. The cost where it does apply is one
+ // 8-byte slot.
+ MachineFrameInfo &GHCMFI = MF.getFrameInfo();
+
+ if (RS && GHCMFI.estimateStackSize(MF) > 0) {
+ const TargetRegisterInfo *TRI = MF.getSubtarget().getRegisterInfo();
+ const TargetRegisterClass &RC = AArch64::GPR64RegClass;
+ int FI = GHCMFI.CreateSpillStackObject(TRI->getSpillSize(RC), TRI->getSpillAlign(RC));
+ RS->addScavengingFrameIndex(FI);
+ LLVM_DEBUG(dbgs() << "GHC function with a frame, allocated fi#" << FI
+ << " as the emergency spill slot.\n");
+ }
+
return;
+ }
const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
+147
View File
@@ -0,0 +1,147 @@
# libarmsx3_lsfg.so -- Lossless Scaling frame generation, sealed away from the emulator core.
#
# The entire reason this is a separate shared object is symbol collision. volk defines 655 globals
# named vkCreateImage, vkQueueSubmit, ... and all 124 that our Vulkan loader declares in
# rpcs3/Emu/RSX/VK/vk_android_loader.h are among them. Linked into libarmsx3-core.so this either
# fails at link or, worse, merges -- and framegen's volkLoadDevice(itsOwnDevice) then repoints the
# whole RSX renderer at framegen's VkDevice. See armsx3_lsfg_shim.h.
#
# Android only. framegen's non-Android path shares images by FD, which Adreno and Mali refuse for
# AHB-imported memory, so there is nothing here worth building for desktop.
if (NOT ANDROID)
return()
endif()
set(LSFG_ROOT "${CMAKE_CURRENT_SOURCE_DIR}/lsfg-vk-android")
if (NOT EXISTS "${LSFG_ROOT}/framegen/CMakeLists.txt")
message(STATUS "LSFG: 3rdparty/lsfg/lsfg-vk-android is missing, frame generation will not be built")
return()
endif()
if (NOT EXISTS "${LSFG_ROOT}/thirdparty/volk/volk.c")
# Called out explicitly because the failure is otherwise mystifying: framegen links volk
# PUBLIC, so without it the error names framegen rather than the submodule that is missing.
message(STATUS "LSFG: thirdparty/volk is missing (git submodule update --init), skipping")
return()
endif()
# volk, built for Android.
#
# VK_USE_PLATFORM_ANDROID_KHR has to be set on VOLK ITSELF, not only on framegen. Without it volk
# never defines vkGetAndroidHardwareBufferPropertiesANDROID, and the resulting undefined symbol
# points at framegen -- sending you to debug the wrong target entirely.
add_library(armsx3_lsfg_volk STATIC "${LSFG_ROOT}/thirdparty/volk/volk.c")
target_include_directories(armsx3_lsfg_volk PUBLIC "${LSFG_ROOT}/thirdparty/volk")
target_compile_definitions(armsx3_lsfg_volk PUBLIC VK_USE_PLATFORM_ANDROID_KHR VK_NO_PROTOTYPES)
set_target_properties(armsx3_lsfg_volk PROPERTIES
POSITION_INDEPENDENT_CODE ON
C_VISIBILITY_PRESET hidden)
# framegen.
#
# Its own CMakeLists expects a target called `volk`, so alias ours rather than patching upstream.
if (NOT TARGET volk)
add_library(volk ALIAS armsx3_lsfg_volk)
endif()
add_subdirectory("${LSFG_ROOT}/framegen" "${CMAKE_CURRENT_BINARY_DIR}/framegen" EXCLUDE_FROM_ALL)
# Undo the project-wide -fno-exceptions for framegen and the shim.
#
# The top-level build sets it with add_compile_options, which every later add_subdirectory
# inherits. framegen has dozens of throw sites and they do not warn -- they fail to compile. The
# shim needs exceptions for the opposite reason: it exists to CATCH them so none reach the dlopen
# boundary.
foreach (tgt lsfg-vk-framegen)
if (TARGET ${tgt})
target_compile_options(${tgt} PRIVATE -fexceptions)
set_target_properties(${tgt} PROPERTIES
POSITION_INDEPENDENT_CODE ON
CXX_VISIBILITY_PRESET hidden
VISIBILITY_INLINES_HIDDEN ON)
# PRIVATE, not PUBLIC: leaking this onto consumers collides with the valueless #define
# our own Vulkan headers use, in hundreds of RSX translation units.
target_compile_definitions(${tgt} PRIVATE VK_USE_PLATFORM_ANDROID_KHR)
endif()
endforeach()
# Shader extraction: DXBC out of the user's own Lossless.dll, translated to SPIR-V.
#
# framegen asks for SPIR-V by name and does not read the DLL itself, so this chain is the caller's
# responsibility. Building upstream's own libraries rather than writing a DXBC translator: dxbc is
# DXVK's, and reimplementing it would be absurd.
#
# Optional. Without these the library still builds and frame generation still reports itself
# available -- it just cannot initialize until shaders exist, which is also what happens when the
# user has not supplied a DLL.
set(LSFG_HAS_EXTRACT OFF)
if (EXISTS "${LSFG_ROOT}/thirdparty/dxbc/CMakeLists.txt" AND
EXISTS "${LSFG_ROOT}/thirdparty/pe-parse/CMakeLists.txt")
# pe-parse defaults to a shared library and command-line tools, neither of which belongs in
# an APK. Forced here because its options are plain option(), so they take whatever is
# already in the cache unless overridden.
set(BUILD_SHARED_LIBS OFF CACHE BOOL "" FORCE)
set(BUILD_COMMAND_LINE_TOOLS OFF CACHE BOOL "" FORCE)
set(PEPARSE_ENABLE_EXAMPLES OFF CACHE BOOL "" FORCE)
set(PEPARSE_ENABLE_TESTING OFF CACHE BOOL "" FORCE)
add_subdirectory("${LSFG_ROOT}/thirdparty/dxbc" "${CMAKE_CURRENT_BINARY_DIR}/dxbc" EXCLUDE_FROM_ALL)
add_subdirectory("${LSFG_ROOT}/thirdparty/pe-parse" "${CMAKE_CURRENT_BINARY_DIR}/pe-parse" EXCLUDE_FROM_ALL)
foreach (tgt dxbc pe-parse)
if (TARGET ${tgt})
# Same -fno-exceptions problem as framegen: both throw, and inheriting the
# project-wide flag turns that into a compile error rather than a warning.
target_compile_options(${tgt} PRIVATE -fexceptions)
set_target_properties(${tgt} PROPERTIES
POSITION_INDEPENDENT_CODE ON
CXX_VISIBILITY_PRESET hidden)
set(LSFG_HAS_EXTRACT ON)
endif()
endforeach()
endif()
add_library(armsx3_lsfg SHARED armsx3_lsfg_shim.cpp)
target_include_directories(armsx3_lsfg PRIVATE
"${CMAKE_CURRENT_SOURCE_DIR}"
"${LSFG_ROOT}/framegen/public")
target_compile_options(armsx3_lsfg PRIVATE -fexceptions)
target_compile_definitions(armsx3_lsfg PRIVATE VK_USE_PLATFORM_ANDROID_KHR)
set_target_properties(armsx3_lsfg PROPERTIES
CXX_STANDARD 20
CXX_STANDARD_REQUIRED ON
CXX_VISIBILITY_PRESET hidden
VISIBILITY_INLINES_HIDDEN ON
OUTPUT_NAME "armsx3_lsfg")
target_link_libraries(armsx3_lsfg PRIVATE lsfg-vk-framegen armsx3_lsfg_volk android log)
if (LSFG_HAS_EXTRACT)
target_sources(armsx3_lsfg PRIVATE
"${LSFG_ROOT}/src/extract/trans.cpp"
"${LSFG_ROOT}/src/extract/extract.cpp")
target_include_directories(armsx3_lsfg PRIVATE "${LSFG_ROOT}/include")
target_link_libraries(armsx3_lsfg PRIVATE dxbc pe-parse)
target_compile_definitions(armsx3_lsfg PRIVATE ARMSX3_LSFG_HAVE_EXTRACT=1)
message(STATUS "LSFG: shader extraction enabled (dxbc + pe-parse)")
else()
message(STATUS "LSFG: shader extraction NOT available, frame generation cannot initialize")
endif()
# Keep the exported surface to the shim alone.
#
# The version script is what makes the isolation real rather than aspirational: without it,
# framegen's and volk's symbols are still dynamic and the loader can bind our renderer's vk* to
# them. Verify with:
# llvm-nm --defined-only --extern-only libarmsx3_lsfg.so
# Only armsx3_lsfg_* may appear. Any vk* or LSFG_3_1 symbol means this stopped working.
target_link_options(armsx3_lsfg PRIVATE
"-Wl,--version-script=${CMAKE_CURRENT_SOURCE_DIR}/armsx3_lsfg.map"
"-Wl,--no-undefined")
+32
View File
@@ -0,0 +1,32 @@
/* Exported surface of libarmsx3_lsfg.so.
*
* This list IS the isolation. framegen and volk are statically linked into this library and
* between them define 655 globals named vkCreateImage, vkQueueSubmit, ... -- 124 of which are
* exactly the names libarmsx3-core.so's Vulkan loader declares. If any of those stay dynamic,
* the loader is free to bind the renderer's entry points to framegen's copies, and framegen's
* volkLoadDevice() has already pointed those at a different VkDevice.
*
* -fvisibility=hidden covers most of it; this covers the rest, including anything upstream marks
* __attribute__((visibility("default"))) -- which framegen's public API does.
*
* Check it, do not assume it:
* llvm-nm --defined-only --extern-only libarmsx3_lsfg.so
* Nothing but armsx3_lsfg_* should be listed.
*/
{
global:
armsx3_lsfg_abi_version;
armsx3_lsfg_initialize;
armsx3_lsfg_create_context_ahb;
armsx3_lsfg_present;
armsx3_lsfg_destroy_context;
armsx3_lsfg_wait_idle;
armsx3_lsfg_finalize;
armsx3_lsfg_last_error;
armsx3_lsfg_import_shaders;
armsx3_lsfg_shader_count;
armsx3_lsfg_get_shader;
local:
*;
};
+404
View File
@@ -0,0 +1,404 @@
// Implementation of the C ABI in armsx3_lsfg_shim.h.
//
// This translation unit is the ONLY thing in libarmsx3_lsfg.so that anyone outside it may touch.
// Everything else -- framegen, volk, and volk's 655 vk* globals -- stays hidden behind
// -fvisibility=hidden so the dynamic linker cannot bind our renderer's vkCmdDraw to framegen's
// copy. See the header for why that matters.
//
// Rules for every entry point here:
// * no C++ type crosses the boundary (separate libc++ per .so under c++_static),
// * no exception crosses the boundary (framegen throws; dlopen'd code must not),
// * a failure returns a code and leaves a message in armsx3_lsfg_last_error().
#include "armsx3_lsfg_shim.h"
#include <lsfg_3_1.hpp>
#include <lsfg_3_1p.hpp>
#ifdef ARMSX3_LSFG_HAVE_EXTRACT
#include <extract/extract.hpp>
#include <extract/trans.hpp>
#include <config/config.hpp>
#endif
#include <exception>
#include <map>
#include <string>
#include <vector>
namespace
{
// thread_local because the renderer and whatever calls initialize() are not the same thread,
// and a shared buffer would let one overwrite the other's message mid-report.
thread_local std::string g_last_error;
bool g_initialized = false;
// Which shader family initialize() chose. Fixed until finalize(): LSFG_3_1 and LSFG_3_1P keep
// entirely separate device state and context tables, so a context created by one cannot be
// presented or destroyed through the other -- every entry point below has to dispatch on this.
bool g_performance = false;
void clear_error()
{
g_last_error.clear();
}
void set_error(const char* what)
{
g_last_error = what ? what : "unknown error";
}
void set_error(const std::string& what)
{
g_last_error = what.empty() ? "unknown error" : what;
}
}
// Wrap a call so nothing escapes.
//
// catch (...) rather than catching LSFG's types: framegen throws several, they are not part of
// its public header, and an exception reaching the dlopen boundary is undefined behaviour -- so
// the exact type matters less than the guarantee that none of them get out.
#define ARMSX3_LSFG_GUARD(expr, failure_result) \
try \
{ \
clear_error(); \
expr; \
} \
catch (const std::exception& e) \
{ \
set_error(e.what()); \
return (failure_result); \
} \
catch (...) \
{ \
set_error("unknown exception from framegen"); \
return (failure_result); \
}
extern "C" uint32_t armsx3_lsfg_abi_version(void)
{
return ARMSX3_LSFG_ABI_VERSION;
}
extern "C" const char* armsx3_lsfg_last_error(void)
{
return g_last_error.c_str();
}
extern "C" int armsx3_lsfg_initialize(uint64_t device_uuid, int is_hdr, float flow_scale,
uint64_t generation_count, int performance, armsx3_lsfg_shader_loader loader, void* user)
{
if (!loader)
{
set_error("no shader loader supplied");
return ARMSX3_LSFG_ERR_BAD_ARGUMENT;
}
// The std::function is built HERE, on framegen's side of the boundary, from a plain C
// function pointer. That is the whole point of taking a function pointer in the header: an
// std::function constructed by the core would be a different type under a different libc++.
//
// Throwing out of this lambda is how a missing shader is reported to framegen, which is what
// it expects -- and the throw stays inside this .so, caught by the guard below.
const auto bridge = [loader, user](const std::string& name) -> std::vector<uint8_t>
{
const uint8_t* data = nullptr;
uint32_t size = 0;
if (loader(name.c_str(), &data, &size, user) != ARMSX3_LSFG_OK || !data || !size)
{
throw std::runtime_error("shader not available: " + name);
}
return std::vector<uint8_t>(data, data + size);
};
// Recorded BEFORE the call so the guard's failure path cannot leave the two disagreeing.
g_performance = performance != 0;
if (g_performance)
{
ARMSX3_LSFG_GUARD(
LSFG_3_1P::initialize(device_uuid, is_hdr != 0, flow_scale, generation_count, bridge),
ARMSX3_LSFG_ERR_SHADERS)
}
else
{
ARMSX3_LSFG_GUARD(
LSFG_3_1::initialize(device_uuid, is_hdr != 0, flow_scale, generation_count, bridge),
ARMSX3_LSFG_ERR_SHADERS)
}
g_initialized = true;
return ARMSX3_LSFG_OK;
}
extern "C" int32_t armsx3_lsfg_create_context_ahb(void* in0, void* in1, void* const* out_n,
uint32_t out_count, uint32_t width, uint32_t height, int32_t format)
{
if (!g_initialized)
{
set_error("not initialized");
return ARMSX3_LSFG_ERR_NOT_INITIALIZED;
}
if (!in0 || !in1 || !out_n || !out_count)
{
set_error("null image or empty output set");
return ARMSX3_LSFG_ERR_BAD_ARGUMENT;
}
int32_t id = ARMSX3_LSFG_ERR_UNKNOWN;
// AHardwareBuffer* arrives as void* so the header stays free of android/hardware_buffer.h,
// which the core has no reason to include.
std::vector<AHardwareBuffer*> outs;
outs.reserve(out_count);
for (uint32_t i = 0; i < out_count; ++i)
{
outs.push_back(static_cast<AHardwareBuffer*>(out_n[i]));
}
if (g_performance)
{
ARMSX3_LSFG_GUARD(
id = LSFG_3_1P::createContextFromAHB(
static_cast<AHardwareBuffer*>(in0), static_cast<AHardwareBuffer*>(in1), outs,
VkExtent2D{width, height}, static_cast<VkFormat>(format)),
ARMSX3_LSFG_ERR_VULKAN)
}
else
{
ARMSX3_LSFG_GUARD(
id = LSFG_3_1::createContextFromAHB(
static_cast<AHardwareBuffer*>(in0), static_cast<AHardwareBuffer*>(in1), outs,
VkExtent2D{width, height}, static_cast<VkFormat>(format)),
ARMSX3_LSFG_ERR_VULKAN)
}
return id;
}
extern "C" int armsx3_lsfg_present(int32_t ctx, int in_sem, const int* out_sems, uint32_t out_count)
{
if (!g_initialized)
{
set_error("not initialized");
return ARMSX3_LSFG_ERR_NOT_INITIALIZED;
}
std::vector<int> outs;
outs.reserve(out_count);
for (uint32_t i = 0; i < out_count; ++i)
{
outs.push_back(out_sems ? out_sems[i] : -1);
}
if (g_performance)
{
ARMSX3_LSFG_GUARD(LSFG_3_1P::presentContext(ctx, in_sem, outs), ARMSX3_LSFG_ERR_VULKAN)
}
else
{
ARMSX3_LSFG_GUARD(LSFG_3_1::presentContext(ctx, in_sem, outs), ARMSX3_LSFG_ERR_VULKAN)
}
return ARMSX3_LSFG_OK;
}
extern "C" int armsx3_lsfg_destroy_context(int32_t ctx)
{
if (!g_initialized)
{
return ARMSX3_LSFG_OK; // nothing to release
}
if (g_performance)
{
ARMSX3_LSFG_GUARD(LSFG_3_1P::deleteContext(ctx), ARMSX3_LSFG_ERR_UNKNOWN)
}
else
{
ARMSX3_LSFG_GUARD(LSFG_3_1::deleteContext(ctx), ARMSX3_LSFG_ERR_UNKNOWN)
}
return ARMSX3_LSFG_OK;
}
extern "C" void armsx3_lsfg_wait_idle(void)
{
if (!g_initialized)
{
return;
}
try
{
if (g_performance) LSFG_3_1P::waitIdle(); else LSFG_3_1::waitIdle();
}
catch (...)
{
// Deliberately swallowed and not recorded: this is called on the present path, and a
// failure to wait is reported by whatever uses the images next. Setting the error string
// here would overwrite a more useful message from the call that actually failed.
}
}
extern "C" void armsx3_lsfg_finalize(void)
{
if (!g_initialized)
{
return;
}
try
{
if (g_performance) LSFG_3_1P::finalize(); else LSFG_3_1::finalize();
}
catch (...)
{
}
g_initialized = false;
}
#ifdef ARMSX3_LSFG_HAVE_EXTRACT
// Satisfy the one symbol upstream's extract.cpp needs from its config layer.
//
// It reads exactly one field, Config::activeConf.dll, to find the file. Defining the object here
// rather than compiling their config module avoids dragging in toml11 and a config-file format
// that has no meaning inside an APK -- the path comes from the user's file picker instead.
namespace Config { Configuration activeConf; }
namespace
{
// name -> SPIR-V, translated once at import.
std::map<std::string, std::vector<uint8_t>> g_shaders;
}
extern "C" int armsx3_lsfg_import_shaders(const char* dll_path)
{
if (!dll_path || !*dll_path)
{
set_error("no file selected");
return ARMSX3_LSFG_ERR_BAD_ARGUMENT;
}
clear_error();
g_shaders.clear();
// Upstream's own shader names, both families.
//
// Taken verbatim from nameIdxTable in extract.cpp rather than guessed -- a made-up name fails
// as "Shader hash not found", which reads like a corrupt DLL and is not.
//
// Two sets: the plain names are LSFG 3.1 and the p_ prefixed ones are 3.1p. Which family gets
// used depends on which framegen entry point runs, so both are extracted and whatever the DLL
// actually contains is kept. Missing names are skipped rather than fatal, because a given
// Lossless Scaling version legitimately ships only one family.
static const char* const k_names[] = {
"mipmaps", "alpha[0]", "alpha[1]", "alpha[2]", "alpha[3]",
"beta[0]", "beta[1]", "beta[2]", "beta[3]", "beta[4]",
"gamma[0]", "gamma[1]", "gamma[2]", "gamma[3]", "gamma[4]",
"delta[0]", "delta[1]", "delta[2]", "delta[3]", "delta[4]",
"delta[5]", "delta[6]", "delta[7]", "delta[8]", "delta[9]",
"generate",
"p_mipmaps", "p_alpha[0]", "p_alpha[1]", "p_alpha[2]", "p_alpha[3]",
"p_beta[0]", "p_beta[1]", "p_beta[2]", "p_beta[3]", "p_beta[4]",
"p_gamma[0]", "p_gamma[1]", "p_gamma[2]", "p_gamma[3]", "p_gamma[4]",
"p_delta[0]", "p_delta[1]", "p_delta[2]", "p_delta[3]", "p_delta[4]",
"p_delta[5]", "p_delta[6]", "p_delta[7]", "p_delta[8]", "p_delta[9]",
"p_generate",
};
try
{
Config::activeConf.dll = dll_path;
Extract::extractShaders();
for (const char* name : k_names)
{
// getShader hands back DXBC; framegen wants SPIR-V. Translating at import rather than
// on demand keeps the cost off the present path entirely.
//
// Individually guarded: a DLL that ships only one shader family throws on every name
// in the other, and that is normal rather than a failure of the import.
try
{
auto spirv = Extract::translateShader(Extract::getShader(name));
if (!spirv.empty())
{
g_shaders[name] = std::move(spirv);
}
}
catch (const std::exception&)
{
// Not in this DLL. Keep going.
}
}
if (g_shaders.empty())
{
set_error("no usable shaders in that file -- is it Lossless.dll from Lossless Scaling?");
return ARMSX3_LSFG_ERR_SHADERS;
}
}
catch (const std::exception& e)
{
g_shaders.clear();
set_error(e.what());
return ARMSX3_LSFG_ERR_SHADERS;
}
catch (...)
{
g_shaders.clear();
set_error("unknown failure reading the file");
return ARMSX3_LSFG_ERR_SHADERS;
}
return static_cast<int>(g_shaders.size());
}
extern "C" int armsx3_lsfg_shader_count(void)
{
return static_cast<int>(g_shaders.size());
}
extern "C" int armsx3_lsfg_get_shader(const char* name, const uint8_t** out_data, uint32_t* out_size)
{
if (!name || !out_data || !out_size)
{
return ARMSX3_LSFG_ERR_BAD_ARGUMENT;
}
const auto it = g_shaders.find(name);
if (it == g_shaders.end() || it->second.empty())
{
return ARMSX3_LSFG_ERR_SHADERS;
}
*out_data = it->second.data();
*out_size = static_cast<uint32_t>(it->second.size());
return ARMSX3_LSFG_OK;
}
#else
extern "C" int armsx3_lsfg_import_shaders(const char*)
{
set_error("this build has no shader extraction support");
return ARMSX3_LSFG_ERR_SHADERS;
}
extern "C" int armsx3_lsfg_shader_count(void) { return 0; }
extern "C" int armsx3_lsfg_get_shader(const char*, const uint8_t**, uint32_t*)
{
return ARMSX3_LSFG_ERR_SHADERS;
}
#endif
+142
View File
@@ -0,0 +1,142 @@
// C ABI for Lossless Scaling frame generation.
//
// framegen CANNOT be linked into libarmsx3-core.so. It links volk, which defines 655 globals
// named vkCreateImage, vkQueueSubmit, ... and 124 of those are byte-for-byte the names our own
// Vulkan loader declares in rpcs3/Emu/RSX/VK/vk_android_loader.h -- every single symbol the RSX
// renderer uses. Two ways that goes wrong, and the second is the one that costs a week:
//
// 1. duplicate symbol at link time (clang defaults to -fno-common), or
// 2. the linker merges them, and framegen's volkLoadDevice(itsOwnDevice) then repoints every
// entry point the renderer uses at framegen's VkDevice. Every later vkCmdDraw goes to the
// wrong device, and it presents as a driver crash with nothing pointing at frame generation.
//
// So framegen and volk live in their own libarmsx3_lsfg.so, reached by dlopen + dlsym through
// this header. Nothing here is C++: the CMake project builds ANDROID_STL=c++_static, so each .so
// carries its own libc++ and an std::vector or std::function crossing the boundary would be two
// unrelated types that happen to share a name. The shim builds those on its own side.
//
// framegen also throws (LSFG::vulkan_error and friends). Exceptions must not cross a dlopen
// boundary either, so every entry point here catches everything and returns a code; the message
// is retrievable with armsx3_lsfg_last_error().
#pragma once
#include <stdint.h>
#ifdef __cplusplus
extern "C" {
#endif
// Bump when anything below changes shape. The loader refuses a library whose version it does not
// recognise, so a stale libarmsx3_lsfg.so on a user's device fails loudly at load instead of
// quietly passing mismatched structs.
#define ARMSX3_LSFG_ABI_VERSION 2u
// Mark the exported surface explicitly.
//
// The library is built -fvisibility=hidden so framegen's and volk's symbols stay in, and a
// version script narrows the dynamic table further. Neither of those can PROMOTE a symbol: a
// function hidden at compile time is local in the object, and `global:` in the linker script
// cannot bring it back. Without this attribute the .so builds and exports nothing at all, and
// the failure only shows up as dlsym returning null at runtime.
#if defined(__GNUC__) || defined(__clang__)
#define ARMSX3_LSFG_API __attribute__((visibility("default")))
#else
#define ARMSX3_LSFG_API
#endif
enum armsx3_lsfg_result
{
ARMSX3_LSFG_OK = 0,
ARMSX3_LSFG_ERR_UNKNOWN = -1,
ARMSX3_LSFG_ERR_NOT_INITIALIZED = -2,
ARMSX3_LSFG_ERR_BAD_ARGUMENT = -3,
ARMSX3_LSFG_ERR_SHADERS = -4,
ARMSX3_LSFG_ERR_VULKAN = -5,
};
// Hand back the SPIR-V for a named shader.
//
// framegen does NOT read Lossless.dll -- it asks for shaders by name and expects SPIR-V back.
// Extracting them from the user's own copy (PE resource -> DXBC -> SPIR-V) is the caller's job,
// which is deliberate: the shaders are THS's property and nothing here ships or downloads them.
//
// Return ARMSX3_LSFG_OK and set *out_data / *out_size on success. The buffer must stay valid
// until the initialize() call that triggered this returns. Any other return means "no such
// shader" and fails initialization.
typedef int (*armsx3_lsfg_shader_loader)(const char* name, const uint8_t** out_data,
uint32_t* out_size, void* user);
// Version of the loaded library. Call first; anything else on a mismatched library is undefined.
ARMSX3_LSFG_API uint32_t armsx3_lsfg_abi_version(void);
// Bring up framegen on the adapter identified by device_uuid (VkPhysicalDeviceIDProperties
// deviceUUID, 16 bytes, passed as the first 8 -- that is what framegen matches on).
//
// framegen creates its OWN VkDevice on that adapter. It does not share ours, which is why images
// have to be handed over as AHardwareBuffer below rather than as VkImage.
// performance selects framegen's 3.1p shader family instead of 3.1: a cheaper pipeline at lower
// quality, which is the difference between usable and not on a mobile GPU. It is fixed for the
// lifetime of the library state -- every context, present and teardown after this call goes to the
// family chosen here, because the two keep separate contexts and separate device state.
//
// flow_scale is the optical-flow resolution as a fraction of full: 1.0 is upstream's default and
// lower is cheaper. Note the sense is inverted from upstream's own config file, which stores a
// divisor and passes 1.0f/value here.
ARMSX3_LSFG_API int armsx3_lsfg_initialize(uint64_t device_uuid, int is_hdr, float flow_scale,
uint64_t generation_count, int performance, armsx3_lsfg_shader_loader loader, void* user);
// Create a context over a set of shared images.
//
// AHardwareBuffer rather than the FD path framegen also offers, because Adreno and Mali both
// refuse vkGetMemoryFdKHR(OPAQUE_FD) on AHB-imported memory -- the FD path simply does not work
// on the hardware this port runs on.
//
// The caller keeps ownership of every AHardwareBuffer and must keep them alive until the context
// is destroyed. Returns a context id >= 0, or a negative armsx3_lsfg_result.
ARMSX3_LSFG_API int32_t armsx3_lsfg_create_context_ahb(void* in0, void* in1, void* const* out_n,
uint32_t out_count, uint32_t width, uint32_t height, int32_t format);
// Generate frames for one presented pair.
//
// Semaphores are sync file descriptors, not VkSemaphore: framegen is on a different device and a
// VkSemaphore handle would be meaningless to it. in_sem is waited on before generation starts;
// each out_sems[i] is signalled when output image i is ready. Pass -1 for an unused slot.
ARMSX3_LSFG_API int armsx3_lsfg_present(int32_t ctx, int in_sem, const int* out_sems, uint32_t out_count);
ARMSX3_LSFG_API int armsx3_lsfg_destroy_context(int32_t ctx);
// Read the user's own Lossless.dll and keep the shaders it contains.
//
// Nothing is bundled or downloaded: the shaders are THS's property and the user must supply a
// legitimately purchased copy. Only the extracted SPIR-V is kept -- the DLL itself is not needed
// afterwards and the caller may delete its copy.
//
// The work is PE resource walk -> DXBC -> SPIR-V, and it is slow enough to be worth doing once
// and caching rather than at every boot. Returns the number of shaders extracted, or a negative
// armsx3_lsfg_result; armsx3_lsfg_last_error() explains a failure in terms a user can act on
// ("is Lossless Scaling up to date?" rather than a resource id).
ARMSX3_LSFG_API int armsx3_lsfg_import_shaders(const char* dll_path);
// How many shaders are currently held. Zero means frame generation cannot start.
ARMSX3_LSFG_API int armsx3_lsfg_shader_count(void);
// Serve a previously imported shader by name, for initialize()'s loader.
//
// Pass a null loader to armsx3_lsfg_initialize to use these instead of supplying your own.
ARMSX3_LSFG_API int armsx3_lsfg_get_shader(const char* name, const uint8_t** out_data, uint32_t* out_size);
// Block until framegen's device is idle.
//
// Needed on Android because framegen's device reads AHBs that OUR device writes, and there is no
// semaphore shared between the two. Without this the read races the write. It is also the reason
// frame generation cannot be free here: this is a device-level stall, not a queue wait.
ARMSX3_LSFG_API void armsx3_lsfg_wait_idle(void);
ARMSX3_LSFG_API void armsx3_lsfg_finalize(void);
// Message for the last failing call on this thread, or "" if none. Never null.
ARMSX3_LSFG_API const char* armsx3_lsfg_last_error(void);
#ifdef __cplusplus
}
#endif
+8
View File
@@ -0,0 +1,8 @@
# Oboe
#
# Android-only. Oboe wraps AAudio (and OpenSL ES on older devices) and carries a
# per-device quirks database plus stream-restart handling, which is the part that
# matters on the low-end parts where plain AAudio glitches.
add_subdirectory(oboe EXCLUDE_FROM_ALL)
add_library(3rdparty::oboe ALIAS oboe)
Vendored Submodule
+1
Submodule 3rdparty/oboe/oboe added at 0da326e4ef
-31
View File
@@ -1,39 +1,8 @@
ARMSX3
======
Proof of concept Android port of RPCS3.
Uses the latest RPCS3 upstream code (the recent ARM64 improvements included).
Status
------
From my testing, I only tried Skate 3. It boots, loads and reaches gameplay at roughly 20 to 30 fps on a
Snapdragon 8 Gen 2. Rendering, audio, touch controls and physical controllers
work. Almost nothing else has been tested. So the main stop gap at the moment is performance/speed.
Differences from upstream RPCS3
-------------------------------
Some of the fixes here are not in upstream and affect any ARM64 build, not only
Android:
* Shaders declared runtime sized arrays inside uniform blocks, which requires
VK_EXT_shader_uniform_buffer_unsized_array. Adreno does not support that
extension, so every game pipeline failed to compile and nothing rendered.
Concrete array bounds are emitted when the extension is missing.
* The ARM64 SPU block verification checksum folded two thirds of every block
through an absolute difference. That collides on the near identical job
binaries an SPU job manager streams through the same local store address, so
a cached block could end up running against another job's code. It sums now.
* Thread affinity was compiled out on Android, and the core had no ARM
big.LITTLE topology, so SPU and RSX threads were never placed on the fast
cores.
* The LLVM JIT target was pinned to cortex-a34, an in order core from 2016. It
detects the host now.
Building
--------
+34 -3
View File
@@ -1,3 +1,4 @@
#include <cerrno>
#include "File.h"
#include "mutex.h"
#include "StrFmt.h"
@@ -684,8 +685,20 @@ namespace fs
u64 result = 0;
// Loop because (huge?) read can be processed partially
while (auto r = ::read(m_fd, buffer, count))
for (;;)
{
const auto r = ::read(m_fd, buffer, count);
// EINTR is benign -- a signal landed mid-syscall -- and must be retried, not treated
// as failure. Android app storage is FUSE-backed, where this genuinely happens, and
// the ensure() below turns it into a process abort. Ported in spirit from
// ouroboros420/rpcsx (92144f094).
if (r < 0 && errno == EINTR)
{
continue;
}
if (!r) break; // EOF
ensure(r > 0); // "file::read"
count -= r;
result += r;
@@ -702,8 +715,17 @@ namespace fs
u64 result = 0;
// For safety; see read()
while (auto r = ::pread(m_fd, buffer, count, offset))
for (;;)
{
const auto r = ::pread(m_fd, buffer, count, offset);
// See read(): retry EINTR rather than aborting.
if (r < 0 && errno == EINTR)
{
continue;
}
if (!r) break; // EOF
ensure(r > 0); // "file::read_at"
count -= r;
offset += r;
@@ -721,8 +743,17 @@ namespace fs
u64 result = 0;
// For safety; see read()
while (auto r = ::write(m_fd, buffer, count))
for (;;)
{
const auto r = ::write(m_fd, buffer, count);
// See read(): retry EINTR rather than aborting.
if (r < 0 && errno == EINTR)
{
continue;
}
if (!r) break;
ensure(r > 0); // "file::write"
count -= r;
result += r;
+39
View File
@@ -618,6 +618,22 @@ std::string jit_compiler::cpu(std::string_view _cpu)
m_cpu = fallback_cpu_detection();
}
#ifdef ARCH_ARM64
// Detection reads the MIDR of whichever core happens to be running, so on big.LITTLE it
// can name a small in-order core. JIT'd code runs on every core, so scheduling for the
// smallest one is the wrong default -- fall back to the same wide out-of-order baseline
// used when detection fails outright. Only the schedule/cost model is affected: the
// instruction set still comes from setMAttrs (HWCAP-gated), so this can never emit an
// illegal instruction. Ported from ouroboros420/rpcsx (cc3a18e29), widened to the
// A5xx little cores that modern SoCs actually ship.
if (m_cpu == "cortex-a34" || m_cpu == "cortex-a35" || m_cpu == "cortex-a53" ||
m_cpu == "cortex-a55" || m_cpu == "cortex-a510" || m_cpu == "cortex-a520")
{
jit_log.notice("CPU detection named a little core ('%s'); using cortex-a78 as the schedule baseline.", m_cpu);
m_cpu = "cortex-a78";
}
#endif
if (m_cpu == "sandybridge" ||
m_cpu == "ivybridge" ||
m_cpu == "haswell" ||
@@ -754,6 +770,29 @@ jit_compiler::jit_compiler(const std::unordered_map<std::string, u64>& _link, st
fmt::throw_exception("LLVM Emergency Exit Invoked: '%s'", out);
}, nullptr);
// A separate handler from the fatal one -- LLVM installs and dispatches the two
// independently. Without this, an allocation failure inside LLVM (SmallVector growth
// while codegenning the enormous PPU symbol-resolver module, say) writes "LLVM ERROR:
// out of memory" to fd 2 -- which goes nowhere in an Android app -- and calls abort():
// a signal-6 death with nothing whatsoever in the log. Route it through the same
// recoverable path as the fatal handler, so a guarded compile survives and anything
// else at least says why it died. Allocating inside a bad-alloc handler is
// best-effort, but the failures here are huge single allocations, so a short log
// string still succeeds. Ported from ouroboros420/rpcsx (39a6a4c36).
llvm::remove_bad_alloc_error_handler();
llvm::install_bad_alloc_error_handler([](void*, const char* msg, bool)
{
const std::string_view out = msg ? msg : "";
if (g_llvm_fatal_message)
{
*g_llvm_fatal_message = out;
thread_ctrl::silent_exit();
}
fmt::throw_exception("LLVM Out Of Memory: '%s'", out);
}, nullptr);
return true;
}();
+66 -2
View File
@@ -2191,7 +2191,32 @@ bool handle_access_violation(u32 addr, bool is_writing, bool is_exec, ucontext_t
if (g_tls_access_violation_recovered != addr)
{
vm_log.notice("\n%s", dump_useful_thread_info());
vm_log.always()("[%s] Access violation %s location 0x%x (%s)", cpu->get_name(), is_writing ? "writing" : "reading", addr, (is_writing && vm::check_addr(addr)) ? "read-only memory" : "unmapped memory");
// Name a guest halt for what it is.
//
// The SPU recompilers implement the HALT family (HGT/HEQ/HLGT and friends) by
// storing to 0xffdead00 on purpose, so the fault handler catches it -- see
// make_halt in SPULLVMRecompiler.cpp and its ASMJIT counterpart. Reported as a
// bare access violation it reads like an emulator crash at a nonsense address,
// and it is neither: those instructions are assertions the GAME compiled into
// its own SPU code, so reaching one means the program checked its state, found
// it wrong, and stopped itself. The interesting question is what fed it bad
// data, which is a completely different investigation from a stray pointer.
//
// The interpreter already says "Halt" here; only the recompiled path was
// silent about it. Hit on Eternal Sonata (BLJS10017), whose TCX_CellSpursKernel0
// halts and takes the game's forward progress with it.
if (addr >= 0xffdead00 && addr < 0xffdeae00)
{
vm_log.always()("[%s] SPU halted itself: the guest executed a HALT instruction"
" (trap store to 0x%x). This is the game's own assertion firing, not a bad"
" pointer -- something upstream handed it state it rejected.",
cpu->get_name(), addr);
}
else
{
vm_log.always()("[%s] Access violation %s location 0x%x (%s)", cpu->get_name(), is_writing ? "writing" : "reading", addr, (is_writing && vm::check_addr(addr)) ? "read-only memory" : "unmapped memory");
}
}
// TODO:
@@ -2544,13 +2569,38 @@ static void signal_handler(int /*sig*/, siginfo_t* info, void* uct) noexcept
const bool is_executing = err & 0x10;
const bool is_writing = err & 0x2;
#elif defined(ARCH_ARM64)
const bool is_executing = uptr(info->si_addr) == uptr(RIP(context));
// Guess, replaced below by the hardware's own answer wherever that is available.
//
// This comparison is a heuristic and it decides something load-bearing: is_executing gates
// EVERY recovery path in this handler, so getting it wrong does not merely mislabel a log
// line, it skips handle_access_violation entirely and kills the thread. A data access whose
// faulting address happens to coincide with the PC is classified as an instruction fetch and
// takes that path, and the guest addresses most likely to collide are exactly the ones our
// own mappings sit at.
bool is_executing = uptr(info->si_addr) == uptr(RIP(context));
#if defined(__linux__) || defined(__APPLE__)
// Current CPU state decoder is reverse-engineered from the linux kernel and may not work on other platforms.
const auto decoded_reason = aarch64::decode_fault_reason(context);
const bool is_writing = (decoded_reason == aarch64::fault_reason::data_write);
// ESR_EL1 says what the fault actually was, so prefer it over the address comparison.
//
// Only when the decode produced something meaningful: it returns 'undefined' when the signal
// frame carries no ESR record, and on that path the guess is still the best available answer.
// data_read/data_write are positive evidence that this is NOT an instruction fetch, which is
// the direction that matters -- it is what lets a genuine access violation reach the recovery
// path instead of terminating the thread.
if (decoded_reason == aarch64::fault_reason::data_read ||
decoded_reason == aarch64::fault_reason::data_write)
{
is_executing = false;
}
else if (decoded_reason == aarch64::fault_reason::instruction_execute)
{
is_executing = true;
}
if (decoded_reason != aarch64::fault_reason::data_write &&
decoded_reason != aarch64::fault_reason::data_read)
{
@@ -2746,6 +2796,20 @@ const bool s_terminate_handler_set = []() -> bool
{
std::set_terminate([]()
{
// Re-entrancy guard. Under memory exhaustion the terminate path itself allocates
// (report_fatal_error formats a message -> operator new; with -fno-exceptions a failed
// allocation calls std::terminate again), which recurses forever and buries the real
// crash under a stack of aborts. If terminate re-enters, hard-stop without allocating
// so there is exactly one clean tombstone.
// Ported from ouroboros420/rpcsx (281654906).
static atomic_t<int> s_terminating{0};
if (s_terminating.exchange(1) != 0)
{
::signal(SIGABRT, SIG_DFL);
std::abort();
}
if (IsDebuggerPresent())
{
logs::listener::sync_all();
+7 -36
View File
@@ -87,42 +87,13 @@ android {
}
}
// ARMSX3: fail the build if the bundled ANGLE libraries are not there.
//
// This check exists because of a specific, expensive bug in ARMSX2: the repo's
// blanket `*.so` gitignore rule swallowed the ANGLE prebuilts, they never made it
// into release staging, the APK shipped without them, and the core fell back to
// the system GLES driver in complete silence. Users reported "ANGLE is broken"
// and there was nothing in any log to contradict them.
//
// jniLibs/.gitignore now un-ignores the two files by name. This task is the
// second lock: packaging an APK that claims to support ANGLE without shipping
// ANGLE is a build error, not a runtime surprise. The core-side counterpart is
// the loud error in gl::es::egl_initialize() when the override library is
// selected but cannot be dlopen'd.
val verifyAngleLibs by tasks.registering {
val angleLibs = listOf("libEGL_angle.so", "libGLESv2_angle.so")
val jniLibDir = file("src/main/jniLibs/arm64-v8a")
doLast {
val missing = angleLibs.filter { !File(jniLibDir, it).isFile }
if (missing.isNotEmpty()) {
throw GradleException(
"ANGLE libraries missing from ${'$'}jniLibDir: ${'$'}{missing.joinToString(", ")}.\n" +
"The OpenGL renderer's ANGLE option cannot work without them and would " +
"silently fall back to the system GLES driver.\n" +
"They are tracked in git - check them out, or remove the ANGLE option."
)
}
angleLibs.forEach {
logger.lifecycle("ANGLE: packaging ${'$'}it (${'$'}{File(jniLibDir, it).length()} bytes)")
}
}
}
tasks.matching { it.name.startsWith("merge") && it.name.endsWith("JniLibFolders") }
.configureEach { dependsOn(verifyAngleLibs) }
// ARMSX3: the ANGLE prebuilts and the verifyAngleLibs task that guarded them used
// to live here. They now live in android/armsx3-ui, which is the module that
// actually ships (applicationId com.armsx3) and the module whose UI exposes the
// OpenGL renderer's ANGLE option. This module builds nothing that ships, so the
// guard here could never protect the APK it was written for -- and it never ran at
// all: its message was written with `${'$'}`-style template escaping, and the `", "`
// inside it closed the Kotlin string early, so this file did not compile.
base.archivesName = "rpcsx"
+51 -3
View File
@@ -27,10 +27,13 @@ android {
defaultConfig {
applicationId = "com.armsx3"
minSdk = 26
// Set per variant by android/build-variants.sh: 33 for the A13 build (NDK 28), 35 for
// the A15 build (NDK 29). The core is compiled against the matching API, so these must
// agree -- an APK that installs below its core's target is a dlopen failure at boot.
minSdk = (project.findProperty("armsx3.minSdk") as String?)?.toInt() ?: 33
targetSdk = 37
versionCode = 8
versionName = "0.4.2"
versionCode = 14
versionName = "0.8"
// ARMSX2's UI reads these. STORAGE_ALL_FILES gates the all-files storage path in
// onboarding; IN_APP_UPDATER gates the in-app GitHub-release updater.
@@ -110,6 +113,51 @@ android {
}
}
// ARMSX3: fail the build if the bundled ANGLE libraries are not there.
//
// This check exists because of a specific, expensive bug in ARMSX2: the repo's
// blanket `*.so` gitignore rule swallowed the ANGLE prebuilts, they never made it
// into release staging, the APK shipped without them, and the core fell back to
// the system GLES driver in complete silence. Users reported "ANGLE is broken"
// and there was nothing in any log to contradict them.
//
// jniLibs/.gitignore un-ignores the two files by name. This task is the second
// lock: packaging an APK that claims to support ANGLE without shipping ANGLE is a
// build error, not a runtime surprise. The core-side counterpart is the loud error
// in gl::es::egl_initialize() when the override library is selected but cannot be
// dlopen'd; the UI-side counterpart is the MISSING_LIBS line that
// MainActivityRuntime.applyAngleEnv logs when the option is on and the .so is not
// in nativeLibraryDir.
//
// The claim being guarded is live in THIS module: RendererBackendSection ->
// AngleDriverSection writes Settings.useAngleOpenGL, and applyAngleEnv turns it
// into ARMSX2_ANGLE_EGL_LIBRARY. (Both the libraries and this task used to sit in
// the stale android/armsx3-app module, which builds nothing that ships -- so the
// guard could not fire for the APK it was meant to protect.)
val verifyAngleLibs by tasks.registering {
val angleLibs = listOf("libEGL_angle.so", "libGLESv2_angle.so")
val jniLibDir = file("src/main/jniLibs/arm64-v8a")
doLast {
val missing = angleLibs.filter { !jniLibDir.resolve(it).isFile }
if (missing.isNotEmpty()) {
throw GradleException(
"ANGLE libraries missing from $jniLibDir: ${missing.joinToString(", ")}.\n" +
"The OpenGL renderer's ANGLE option cannot work without them and would " +
"silently fall back to the system GLES driver.\n" +
"They are tracked in git - check them out, or remove the ANGLE option."
)
}
angleLibs.forEach {
logger.lifecycle("ANGLE: packaging $it (${jniLibDir.resolve(it).length()} bytes)")
}
}
}
tasks.matching { it.name.startsWith("merge") && it.name.endsWith("JniLibFolders") }
.configureEach { dependsOn(verifyAngleLibs) }
dependencies {
// Discord Social SDK, staged locally rather than pulled from a repo: it is
// proprietary and distributed per-application from the developer portal.
Binary file not shown.

After

Width:  |  Height:  |  Size: 1.3 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.2 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.4 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.3 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.8 KiB

Some files were not shown because too many files have changed in this diff Show More