mVUupdateFlags was a 1:1 port of x86's double-MOVMSKPS extraction: two NEON
AND/ADDV/UMOV chains glued together with GPR AND/SHL/AND/OR, plus a per-site
weight vector emitted through vixl's literal pool. That is 16 instructions of
flag packing around the 2 instructions of real FMAC work, and it measured as
~29% of all VU1 code time in God of War II - the first VU-bound title we have
profiled. One VU1 program carried 1666 literal-pool slots holding 34 distinct
values, 12.8% of its host code, each on its own cache line.
SLI collapses it. CMLT and FCMEQ both yield all-ones-or-zero lanes, so
sli vZero.4s, vSign.4s, #4
leaves sign in bits [31:4] and zero in [3:0] of one register, and a single AND
against a combined per-lane weight selects lane i's sign into bit (i+4) and its
zero into bit i. The per-lane bit sets are disjoint, so ADDV's sum is an OR.
Seven instructions replace sixteen, in the two temporaries the caller already
had - the weight vector may reuse the sign register, which is dead by then.
Because the weight vector is picked at emit time it also absorbs AND_XYZW and
SHIFT_XYZW, which stop existing as instructions, and the vectors move out of
the literal pool into mVUglob.macWeights so they ride the pinned gprMVUglob
base as a single Ldr and share a few D-cache lines. The status merge folds the
same way: x86's `AND mReg, 0xFF` is provably a no-op off the overflow path, and
the non-sticky copy becomes an ORR shifted operand.
Same rewrite for the COP2 macro path in cop2EmitFlagUpdate, which loads its
weights as a literal - that path deliberately leaves gprMVUglob unpinned, since
x25 is the EE recompiler's RECCYCLE inside an EE block. mVUupdateFlags now
asserts it is micro-mode only. The parked-result copy in q28 goes away too: the
pack only touches q29/q31, so a result in RQSCRATCH survives untouched.
Measured with pcsx2-vurunner on M2 over 323 VU programs: host instructions
-22.8%, cycles -20.0% (median per-program -17.1%, best -34.3%, no program
regressed by more than 0.5%). Emitted host bytes -19.6% to -40.3% across
Burnout 3, R&C UYA, Shadow of the Colossus and Katamari - the I-cache side
should matter more on the SD865's 64 KB L1I than it does here. All 5818 corpus
captures replay bit-identically against the interpreter, before and after.
vu_mac_flag_pack_tests pins the weight table against the interpreter across the
three shapes that select different table rows (full mask, partial mask, and the
single-scalar rotate that is the only folded non-zero shift).
kMvuCompilerAbiVersion 15 -> 16: every flag-writing FMAC changes shape, and the
emitted [x25, #imm] weight offsets only exist in an mVUglob layout carrying
macWeights, so on-disk program caches must evict.
ARMSX2 — Native ARM64 JIT Fork of PCSX2
ARMSX2 is a free and open-source PlayStation 2 (PS2) emulator based on PCSX2. Its purpose is to emulate the PS2's hardware, using a combination of MIPS CPU Interpreters, Recompilers and a Virtual Machine which manages hardware states and PS2 system memory. This allows you to play PS2 games on your phone, PC, or gaming handheld, with many additional features and benefits.
Thank You
The ARMSX2 team is eternally indebted to the PCSX2 project it is based on. We are so fortunate to build on their 20 years of hardcore development.
About This Fork
The upstream PCSX2 project ships an ARM64 interpreter build for ARM, but its high-performance JIT recompilers (EE, IOP, VU0, VU1, and vtlb fast memory) are x86-64 only.
This fork exists to close that gap. The goal is to preserve the correctness features of 20 years of PCSX2 development, while generating the fastest native ARM performance possible.
Current status:
- ✅ EE (Emotion Engine) recompiler — integer, float, MMI, COP0/COP1/COP2, branches, load/store
- ✅ IOP (I/O Processor / R3000A) recompiler — full integer, load/store, branches, coprocessors
- ✅ VU (Vector Unit) recompiler — microVU skeleton + Upper FMAC vector ISA complete; Lower ISA and runtime complete
- ✅ vtlb fast memory
- ✅ Native ARM64 binary builds and boots the PS2 BIOS
- ✅ 2D games are already playable
- ✅ 3D games run
Why LLMs / AI Were Used
A word on methodology:
The x86-64 JIT code in upstream ARMSX2 is already proven correct — it has run thousands of PS2 titles for years. The challenge in this port is not emulator design or JIT theory; it is mechanical translation of a large, well-understood x86-64 assembly codebase into equivalent ARM64 assembly (via VIXL) while preserving the exact same register-allocation contracts, block lifecycle, and recompiler semantics.
Large language models (LLMs) were used as an accelerant for this translation work — pattern-matching x86 JIT boilerplate to ARM64 equivalents, scaffolding emit routines, and keeping the porting velocity high. The JIT logic (block compiler, dispatcher, analysis passes, flag pipelines, clamping rules, Tri-Ace hacks, etc.) is taken directly from the upstream x86 implementation and validated against it. Nothing was hallucinated from scratch.
In other words: the hard engineering was done by the PCSX2 team over two decades. The hard typing — translating ~50k lines of x86 emitter code into ARM64 — is what AI helped compress.
System Requirements
ARMSX2 targets ARM64 across desktop (macOS, Windows, Linux) and mobile (Android, iOS/iPadOS), all from the single shared core. Our setup documentation page contains additional details on software and hardware requirements.
Please note that a BIOS dump from a legitimately-owned PS2 console is required to use the emulator. For more information, visit this page.
Building
Check out our github actions for the latest build recipe
