A masked unpack does not store its quadword with one instruction. doMaskWrite picks, from a sixteen-way switch, a hand-written sequence touching only the lanes that cycle actually writes, and those sequences differ in kind rather than just in offset: a 64-bit store for X+Y, a 64-bit lane store for Z+W, per-lane stores at hand-computed byte offsets for the scattered subsets, and a post-indexed pair for Y+Z. Each is its own chance to name the wrong lane. Only the three-lane subsets were reached. Measured, not assumed: of the sixteen cases, 7/11/13/14 executed and the other twelve had zero counts, because the existing mixed-mask cases happen to protect exactly one lane apiece. The subset is selected by which lanes carry the write-protect code, so ten new cases -- one per unreached subset -- name three protected lanes to reach a single-lane store and two to reach a pair. Protected lanes must come back holding the fill pattern while written lanes hold unpacked data, so a sequence that stores to a neighbouring lane fails on both halves at once. Two more cross the selector with a mode, where the mode merge runs on a partial lane set rather than the whole register. Validated by mutation, each bounded to exactly the predicted set: swapping the Z lane for W in the Y+Z sequence fails write_yz and write_yz_mode1 and nothing else; moving the single-lane Z store from offset 8 to 4 fails write_z alone. The remaining two switch arms stay unreached and are unreachable, which the new absolute test pins from the other side. A fully write-protected cycle is dropped by ProcessMasks before any store is emitted, so the "no lanes" arm is guarded, not exercised; the differential case for it would pass whatever the generator did, since it only has to agree with an oracle that also writes nothing. FullyProtectedBlockWritesNothing asserts the fact itself -- VU memory byte-identical to the fill pattern. The all-lanes arm is likewise dead: the caller emits a plain full-width store when no lane is protected. 1705 tests, 1703 pass, 2 pre-existing skips.
ARMSX2 — Native ARM64 JIT Fork of PCSX2
ARMSX2 is a free and open-source PlayStation 2 (PS2) emulator based on PCSX2. Its purpose is to emulate the PS2's hardware, using a combination of MIPS CPU Interpreters, Recompilers and a Virtual Machine which manages hardware states and PS2 system memory. This allows you to play PS2 games on your phone, PC, or gaming handheld, with many additional features and benefits.
Thank You
The ARMSX2 team is eternally indebted to the PCSX2 project it is based on. We are so fortunate to build on their 20 years of hardcore development.
About This Fork
The upstream PCSX2 project ships an ARM64 interpreter build for ARM, but its high-performance JIT recompilers (EE, IOP, VU0, VU1, and vtlb fast memory) are x86-64 only.
This fork exists to close that gap. The goal is to preserve the correctness features of 20 years of PCSX2 development, while generating the fastest native ARM performance possible.
Current status:
- ✅ EE (Emotion Engine) recompiler — integer, float, MMI, COP0/COP1/COP2, branches, load/store
- ✅ IOP (I/O Processor / R3000A) recompiler — full integer, load/store, branches, coprocessors
- ✅ VU (Vector Unit) recompiler — microVU skeleton + Upper FMAC vector ISA complete; Lower ISA and runtime complete
- ✅ vtlb fast memory
- ✅ Native ARM64 binary builds and boots the PS2 BIOS
- ✅ 2D games are already playable
- ✅ 3D games run
Why LLMs / AI Were Used
A word on methodology:
The x86-64 JIT code in upstream ARMSX2 is already proven correct — it has run thousands of PS2 titles for years. The challenge in this port is not emulator design or JIT theory; it is mechanical translation of a large, well-understood x86-64 assembly codebase into equivalent ARM64 assembly (via VIXL) while preserving the exact same register-allocation contracts, block lifecycle, and recompiler semantics.
Large language models (LLMs) were used as an accelerant for this translation work — pattern-matching x86 JIT boilerplate to ARM64 equivalents, scaffolding emit routines, and keeping the porting velocity high. The JIT logic (block compiler, dispatcher, analysis passes, flag pipelines, clamping rules, Tri-Ace hacks, etc.) is taken directly from the upstream x86 implementation and validated against it. Nothing was hallucinated from scratch.
In other words: the hard engineering was done by the PCSX2 team over two decades. The hard typing — translating ~50k lines of x86 emitter code into ARM64 — is what AI helped compress.
System Requirements
ARMSX2 targets ARM64 across desktop (macOS, Windows, Linux) and mobile (Android, iOS/iPadOS), all from the single shared core. Our setup documentation page contains additional details on software and hardware requirements.
Please note that a BIOS dump from a legitimately-owned PS2 console is required to use the emulator. For more information, visit this page.
Building
Check out our github actions for the latest build recipe
