mirror of
https://github.com/ARMSX2/ARMSX2.git
synced 2026-08-24 16:50:16 -07:00
Boot crash on iOS 18 under LiveContainer (Legacy W^X mode): CPU thread, EXC_BAD_ACCESS KERN_PROTECTION_FAILURE in Arm64BaseBlocks::New -- a str of a branch encoding into an r-x page. In Legacy mode armStartBlock flips only the current block's 1 MiB window to RW, but New()'s link-repoint loop, Remove()'s entry stubs and the backedge patch write into EARLIER blocks, whose windows are execute-protected by then. PatchWord trusted its callers to hold a Begin/EndCodeWrite scope; no caller on those paths did. The toggle modes (macOS, Simulator) masked it because the compile thread's write-protect is off arena-wide during recRecompile. Give the patch primitive its own scope instead: armPatchCodeWord stores through the dual-map alias and, when the target page is outside the open emit window, wraps the store in a page-granular Begin/EndCodeWriteRange. The page check matters -- the bump allocator routinely puts the previous block's link site on the same page as the current block's start, and RX-flipping that page mid-emit would kill the compile in a new way. The whole-arena Begin/EndCodeWrite scopes that only existed to cover those patches are gone with it. They were their own crash: Legacy range windows bypass the refcount, so recClear's paired EndCodeWrite RX-flipped the open emit window whenever the stale-overlap walk cleared blocks mid-compile, and recClearIOP paid a whole-arena mprotect pair per covered IOP store. The fastmem backpatch and the cold-island patches route through armPatchCodeWord for the same reason, and microVU's code-cache scope narrows to its own buffer span so an MTVU close can no longer RX-flip an EE emit window open on the other thread.
449 lines
16 KiB
C++
449 lines
16 KiB
C++
// SPDX-FileCopyrightText: 2026 yaps2 Dev Team
|
|
// SPDX-License-Identifier: GPL-3.0+
|
|
|
|
// ARM64-specific BaseBlocks with block linking that stays signal-safe.
|
|
//
|
|
// x86 BaseBlocks (pcsx2/x86/BaseblockEx.h) uses a std::multimap to track
|
|
// pending link sites and rewrites them whenever a block is added or
|
|
// removed. The multimap is not used here because Remove() runs from the
|
|
// SIGSEGV fastmem handler, where mutating an STL container is unsafe.
|
|
//
|
|
// This implementation supports block linking while keeping Remove() signal-safe:
|
|
// - Link()/New() touch the multimap, but they only run from the
|
|
// compile path (single-threaded, never from a signal).
|
|
// - Remove() does NOT walk the link map. Instead it overwrites the
|
|
// first 4 bytes of each removed block with `B DispatcherReg`, so any
|
|
// stale link re-dispatches through the recLUT. Block memory isn't
|
|
// reclaimed until a full reset, so this 4-byte rewrite always lands
|
|
// on memory the recompiler still owns.
|
|
//
|
|
// Stale sites divert to the DISPATCHER, never to JITCompile. Every link
|
|
// tail stores the target guest pc into cpuRegs.pc before its branch (both
|
|
// recs, all forms — see SetBranchImm / psxSetBranchImm), so a diverted
|
|
// site re-dispatches from a pc that is already correct; the recLUT (the
|
|
// single SMC-invalidation rewrite point) then routes it to the current
|
|
// code, or through JITCompile if the target is genuinely uncompiled.
|
|
// Routing stale sites at JITCompile directly instead re-enters
|
|
// recRecompile on an already-compiled block whenever the target has been
|
|
// removed-and-recompiled — tripping the fnptr assert on Devel, and on
|
|
// Release spuriously superseding live code in place, which strands
|
|
// orphaned callers in the stale pre-SMC compile (the NFS Carbon
|
|
// SLUS-21493 block-cache corruption). AetherSX2's Remove() and the
|
|
// pre-transplant linker enforce the same fallback-to-dispatcher policy.
|
|
//
|
|
// The patch site for each link is the address of a single B instruction
|
|
// emitted by SetBranchImm (see iR5900-arm64.cpp). Aligned 32-bit stores
|
|
// are atomic on AArch64, and `cacheflush(2)` on the patch site is
|
|
// async-signal-safe, so the redirect-stub Remove() write is signal-safe.
|
|
//
|
|
// Range: B imm26 covers ±128 MB. The EE recompiler region is 64 MB
|
|
// (HostMemoryMap::EErecSize), and JITCompile lives in the same region,
|
|
// so all link sites are reachable with a single B.
|
|
//
|
|
// Link-map liveness (2026-07-20). Entries used to accumulate for the whole
|
|
// lifetime of the map: Link() inserts unconditionally, and nothing pruned.
|
|
// Every recompile of a caller emits its branch at a *fresh* code address and
|
|
// appended another entry, so New(pc) re-patched every site the map had ever
|
|
// seen for that pc — each with its own 4-byte icache flush. Measured on
|
|
// Dirge of Cerberus: 13,372 stale sites on one PC, and 93.5% of EE-thread
|
|
// cycles inside __aarch64_sync_cache_range. Each entry now carries the
|
|
// identity of the block that emitted it (owner startpc + that block's code
|
|
// address), and both Link() and New() erase entries whose owner is gone or has
|
|
// since been recompiled. Cost becomes O(live callers) instead of O(all callers
|
|
// ever), and the map stops growing without bound.
|
|
//
|
|
// Liveness rests on two invariants: block code is bump-allocated and never
|
|
// reclaimed until a full Reset() (which also clears the map), so an fnptr
|
|
// value never identifies two different compiles; and BASEBLOCKEX::fnptr is
|
|
// authoritative for the block's *current* code, which is why New() updates
|
|
// it when a startpc is recompiled in place rather than leaving it stale.
|
|
|
|
#pragma once
|
|
|
|
#include <algorithm>
|
|
#include <map>
|
|
|
|
#include "common/HostSys.h"
|
|
#include "arm64/AsmHelpers.h" // armPatchCodeWord (self-scoped W^X code patching)
|
|
#include "x86/BaseblockEx.h" // BASEBLOCK, BASEBLOCKEX, BaseBlockArray, recLUT_SetPage
|
|
|
|
class Arm64BaseBlocks
|
|
{
|
|
protected:
|
|
// A registered branch site, tagged with the block that emitted it so
|
|
// New() can tell live sites from ones left behind by a superseded
|
|
// compile. See the liveness note in the file header.
|
|
struct LinkSite
|
|
{
|
|
// Patch site address, with kLinkSiteCallBit in bit 0.
|
|
uptr site;
|
|
// Block that emitted this site. owner_fnptr == 0 means "untracked":
|
|
// the site was registered outside a compile (only the direct unit
|
|
// tests do this), and is treated as permanently live.
|
|
u32 owner_startpc;
|
|
uptr owner_fnptr;
|
|
};
|
|
|
|
using linkmap_t = std::multimap<u32, LinkSite>;
|
|
|
|
BaseBlockArray blocks;
|
|
linkmap_t links;
|
|
uptr jitcompile = 0;
|
|
uptr dispatcher = 0;
|
|
|
|
// Where a stale site (unresolved Link(), Remove() entry stamp) diverts.
|
|
// The dispatcher when one is registered — re-dispatch from the pc the
|
|
// site's own tail stored — with a JITCompile fallback for direct
|
|
// instantiation (unit tests) that never registers one.
|
|
__fi uptr StaleDispatchTarget() const
|
|
{
|
|
return dispatcher ? dispatcher : jitcompile;
|
|
}
|
|
|
|
// The block currently being compiled, published by New(). Link() stamps
|
|
// it onto every site it records.
|
|
u32 emitting_startpc = 0;
|
|
uptr emitting_fnptr = 0;
|
|
|
|
// Link sites are 4-byte-aligned instruction addresses, so bit 0 of the
|
|
// stored uptr is free: it tags the site's branch form. Untagged = B,
|
|
// tagged = BL (call-ret stack call sites — the BL pushes the hardware
|
|
// return-address stack so the callee's paired RET predicts).
|
|
static constexpr uptr kLinkSiteCallBit = 1;
|
|
|
|
// Encode a B/BL imm26 from `site` to `target`. ARM64 encoding:
|
|
// B : bits[31:26] = 000101, BL: bits[31:26] = 100101
|
|
// bits[25:0] = sign-extended ((target - site) >> 2)
|
|
static u32 EncodeB(uptr site, uptr target, bool call = false)
|
|
{
|
|
const intptr_t off = static_cast<intptr_t>(target) - static_cast<intptr_t>(site);
|
|
pxAssertRel((off & 3) == 0, "Branch offset not 4-byte aligned");
|
|
const intptr_t imm26 = off >> 2;
|
|
pxAssertRel(imm26 >= -(1 << 25) && imm26 < (1 << 25), "Branch offset out of B imm26 range");
|
|
return (call ? 0x94000000u : 0x14000000u) | (static_cast<u32>(imm26) & 0x03FFFFFFu);
|
|
}
|
|
|
|
// Store only, no cache maintenance. Fine for a site inside the block
|
|
// currently being emitted (that buffer has not been executed yet, and
|
|
// armEndBlock() issues one whole-range flush over it), and for the
|
|
// New() repoint loop, whose callers flush via FlushPatchedSites.
|
|
// Never use this on code that is already live without a flush; use
|
|
// PatchAtomic.
|
|
static void PatchWord(uptr site, u32 instr)
|
|
{
|
|
// armPatchCodeWord owns the W^X handling — sites outside the open
|
|
// emit window get their own page-granular write scope. Nothing on
|
|
// the New()/Remove() paths holds one.
|
|
armPatchCodeWord(reinterpret_cast<u8*>(site), instr);
|
|
}
|
|
|
|
// (HostSys::FlushInstructionCache, not the raw builtin: on Darwin the
|
|
// builtin lowers to a compiler-rt ___clear_cache call that the iOS link
|
|
// doesn't provide; the wrapper uses sys_icache_invalidate there.)
|
|
static void FlushRange(uptr lo, uptr hi)
|
|
{
|
|
HostSys::FlushInstructionCache(reinterpret_cast<void*>(lo), static_cast<u32>(hi - lo));
|
|
}
|
|
|
|
static void PatchAtomic(uptr site, u32 instr)
|
|
{
|
|
PatchWord(site, instr);
|
|
// Then make sure cores fetching instructions see the new word.
|
|
FlushRange(site, site + 4);
|
|
// Signal-safe (counter-only) — Remove() can call this from the
|
|
// SIGSEGV fastmem handler.
|
|
}
|
|
|
|
// Drop every entry for `dest_pc` whose owner is gone or has been
|
|
// recompiled since. Called from both Link() and New(): New() alone is not
|
|
// enough, because a destination that is never recompiled again would never
|
|
// have its list revisited, and the entries its callers keep re-registering
|
|
// would accumulate for the life of the map. Erasing a multimap element
|
|
// invalidates only that element's iterator, so the end of the range stays
|
|
// valid across the walk.
|
|
void PruneDeadLinks(u32 dest_pc)
|
|
{
|
|
const auto range = links.equal_range(dest_pc);
|
|
for (auto it = range.first; it != range.second;)
|
|
{
|
|
if (IsOwnerLive(it->second))
|
|
++it;
|
|
else
|
|
it = links.erase(it);
|
|
}
|
|
}
|
|
|
|
// True when `ls` still lives inside a block that is both present and has
|
|
// not been recompiled since the site was registered.
|
|
bool IsOwnerLive(const LinkSite& ls)
|
|
{
|
|
if (ls.owner_fnptr == 0)
|
|
return true; // untracked registration (direct unit tests)
|
|
|
|
const BASEBLOCKEX* owner = Get(ls.owner_startpc);
|
|
return owner && owner->startpc == ls.owner_startpc && owner->fnptr == ls.owner_fnptr;
|
|
}
|
|
|
|
// Coalesce the just-patched sites into as few cache-maintenance ranges as
|
|
// possible. The per-call cost is dominated by the `dsb ish` / `isb` pair,
|
|
// not by the DC/IC ops, so N adjacent 4-byte flushes cost ~N barrier pairs
|
|
// where one range flush costs one. Sites emitted by the same block sit a
|
|
// few instructions apart, so this usually collapses to a single range.
|
|
static void FlushPatchedSites(uptr* sites, u32 count)
|
|
{
|
|
if (count == 0)
|
|
return;
|
|
|
|
std::sort(sites, sites + count);
|
|
|
|
constexpr uptr kMaxGap = 256;
|
|
uptr lo = sites[0], hi = sites[0] + 4;
|
|
for (u32 i = 1; i < count; i++)
|
|
{
|
|
if (sites[i] - hi <= kMaxGap)
|
|
{
|
|
hi = sites[i] + 4;
|
|
continue;
|
|
}
|
|
FlushRange(lo, hi);
|
|
lo = sites[i];
|
|
hi = sites[i] + 4;
|
|
}
|
|
FlushRange(lo, hi);
|
|
}
|
|
|
|
public:
|
|
Arm64BaseBlocks()
|
|
: blocks(0x4000)
|
|
{
|
|
}
|
|
|
|
void SetJITCompile(const void* recompiler_)
|
|
{
|
|
jitcompile = reinterpret_cast<uptr>(recompiler_);
|
|
}
|
|
|
|
void SetDispatcher(const void* dispatcher_)
|
|
{
|
|
dispatcher = reinterpret_cast<uptr>(dispatcher_);
|
|
}
|
|
|
|
// Register a link site that wants to branch directly to the block at
|
|
// `pc`. Patches immediately if a block already exists; otherwise
|
|
// records the site so New(pc, ...) can patch it later. `call` sites
|
|
// get BL instead of B (and keep it across every re-patch).
|
|
void Link(u32 pc, void* patch_site, bool call = false)
|
|
{
|
|
pxAssertRel(jitcompile, "SetJITCompile() must be called before Link()");
|
|
pxAssertRel((reinterpret_cast<uptr>(patch_site) & kLinkSiteCallBit) == 0,
|
|
"Link patch site must be 4-byte aligned");
|
|
|
|
BASEBLOCKEX* target = Get(pc);
|
|
const uptr target_addr = (target && target->startpc == pc)
|
|
? target->fnptr : StaleDispatchTarget();
|
|
// No flush: the site is in the block being emitted right now, and
|
|
// armEndBlock() range-flushes that buffer. See PatchWord.
|
|
PatchWord(reinterpret_cast<uptr>(patch_site),
|
|
EncodeB(reinterpret_cast<uptr>(patch_site), target_addr, call));
|
|
|
|
// Reap this destination's dead entries before adding ours, so its list
|
|
// tracks live callers rather than every caller it has ever had.
|
|
PruneDeadLinks(pc);
|
|
|
|
links.insert({pc,
|
|
LinkSite{reinterpret_cast<uptr>(patch_site) | (call ? kLinkSiteCallBit : 0),
|
|
emitting_startpc, emitting_fnptr}});
|
|
}
|
|
|
|
// Begin compiling the block at `startpc`, whose code starts at `fnptr`.
|
|
// Returns its BASEBLOCKEX, creating one if this startpc has no live block.
|
|
//
|
|
// A startpc can be recompiled while its BASEBLOCKEX is still in the array
|
|
// (recClear resets BLOCK->fnptr across a straddled extent but deliberately
|
|
// spares the in-progress block's entry). BaseBlockArray::insert() has no
|
|
// dedup, so that case must reuse the existing entry — and it must retarget
|
|
// it at the new code, otherwise fnptr keeps pointing at the superseded
|
|
// compile and every consumer of it is wrong: x86size is computed from a
|
|
// dead base, Remove() writes its redirect stub over dead code instead of
|
|
// the live entry, and Link() sends new callers into stale code.
|
|
BASEBLOCKEX* New(u32 startpc, uptr fnptr)
|
|
{
|
|
// Published for Link() to stamp onto the sites this block emits.
|
|
emitting_startpc = startpc;
|
|
emitting_fnptr = fnptr;
|
|
|
|
const int idx = Index(startpc);
|
|
BASEBLOCKEX* block = (idx >= 0 && blocks[idx].startpc == startpc)
|
|
? &blocks[idx] : nullptr;
|
|
if (block)
|
|
block->fnptr = fnptr;
|
|
else
|
|
block = blocks.insert(startpc, fnptr);
|
|
|
|
// Repoint the links waiting on this PC at the new code, dropping any
|
|
// left behind by a block that has since been removed or recompiled.
|
|
// Those sites are unreachable code; patching them is pure waste, and
|
|
// it is that waste which used to dominate the EE thread.
|
|
uptr patched[64];
|
|
u32 npatched = 0;
|
|
|
|
const auto range = links.equal_range(startpc);
|
|
for (auto it = range.first; it != range.second;)
|
|
{
|
|
if (!IsOwnerLive(it->second))
|
|
{
|
|
it = links.erase(it);
|
|
continue;
|
|
}
|
|
|
|
const uptr site = it->second.site & ~kLinkSiteCallBit;
|
|
PatchWord(site, EncodeB(site, fnptr, (it->second.site & kLinkSiteCallBit) != 0));
|
|
if (npatched < std::size(patched))
|
|
patched[npatched++] = site;
|
|
else
|
|
FlushRange(site, site + 4); // overflow: flush as we go
|
|
++it;
|
|
}
|
|
|
|
FlushPatchedSites(patched, npatched);
|
|
return block;
|
|
}
|
|
|
|
int LastIndex(u32 startpc) const
|
|
{
|
|
if (blocks.size() == 0)
|
|
return -1;
|
|
|
|
int imin = 0, imax = (int)blocks.size() - 1, imid;
|
|
|
|
while (imin != imax)
|
|
{
|
|
imid = (imin + imax + 1) >> 1;
|
|
|
|
if (blocks[imid].startpc > startpc)
|
|
imax = imid - 1;
|
|
else
|
|
imin = imid;
|
|
}
|
|
|
|
if (IsDevBuild)
|
|
{
|
|
if (imin != 0)
|
|
pxAssert(blocks[imin].startpc <= startpc);
|
|
if (imin < (int)blocks.size() - 1)
|
|
pxAssert(blocks[imin + 1].startpc > startpc);
|
|
}
|
|
|
|
return imin;
|
|
}
|
|
|
|
__fi int Index(u32 startpc) const
|
|
{
|
|
int idx = LastIndex(startpc);
|
|
|
|
if ((idx == -1) || (startpc < blocks[idx].startpc) ||
|
|
((blocks[idx].size) && (startpc >= blocks[idx].startpc + blocks[idx].size * 4)))
|
|
return -1;
|
|
else
|
|
return idx;
|
|
}
|
|
|
|
__fi BASEBLOCKEX* operator[](int idx)
|
|
{
|
|
if (idx < 0 || idx >= (int)blocks.size())
|
|
return 0;
|
|
|
|
return &blocks[idx];
|
|
}
|
|
|
|
__fi BASEBLOCKEX* Get(u32 startpc)
|
|
{
|
|
return (*this)[Index(startpc)];
|
|
}
|
|
|
|
// Signal-safe: writes a redirect stub at each removed block's entry
|
|
// point so any stale link re-dispatches through the recLUT (see the
|
|
// stale-dispatch policy note in the file header — the diverting site's
|
|
// tail already stored the correct pc, so DispatcherReg routes it to the
|
|
// current code; JITCompile here would re-enter recRecompile on a
|
|
// removed-and-recompiled block), then erases from the flat sorted
|
|
// array. Does NOT touch the link map — mutating an STL container here
|
|
// is not signal-safe. The entries this strands are reaped by the next
|
|
// New() for their destination PC, which sees the owner is gone; until
|
|
// then the redirect stub keeps any site still pointing at this block
|
|
// correct.
|
|
//
|
|
// SL-1: a resident self-loop's back-edge is an internal B to the loop-top
|
|
// (past the entry redirect), so it gets its own atomic repoint — to the
|
|
// block's cold spill stub, which writes back the pinned set, publishes pc,
|
|
// and exits via DispatcherEvent. Flat-array field reads + PatchAtomic
|
|
// only, so the signal-safety contract holds.
|
|
__fi void Remove(int first, int last)
|
|
{
|
|
pxAssert(first <= last);
|
|
|
|
if (jitcompile)
|
|
{
|
|
for (int i = first; i <= last; ++i)
|
|
{
|
|
const uptr site = blocks[i].fnptr;
|
|
PatchAtomic(site, EncodeB(site, StaleDispatchTarget()));
|
|
|
|
if (blocks[i].backedge_site)
|
|
PatchAtomic(blocks[i].backedge_site,
|
|
EncodeB(blocks[i].backedge_site, blocks[i].backedge_stub));
|
|
}
|
|
}
|
|
|
|
blocks.erase(first, last + 1);
|
|
}
|
|
|
|
__fi void Reset()
|
|
{
|
|
blocks.clear();
|
|
links.clear();
|
|
emitting_startpc = 0;
|
|
emitting_fnptr = 0;
|
|
}
|
|
|
|
#ifdef PCSX2_RECOMPILER_TESTS
|
|
// Test-only introspection. Returns true iff a link patch site within the
|
|
// block containing src_pc targets a block at dst_pc. The link multimap
|
|
// is keyed by destination PC, so the entries for dst_pc are walked to check
|
|
// whether the patch site lies inside [block.fnptr, block.fnptr + x86size),
|
|
// where x86size is BASEBLOCKEX's (legacy-named) host machine-code byte size.
|
|
// O(L_d + log B) where L_d is the number of links to dst_pc and B is the
|
|
// block count.
|
|
bool IsLinked(u32 src_pc, u32 dst_pc) const
|
|
{
|
|
const int idx = LastIndex(src_pc);
|
|
if (idx < 0)
|
|
return false;
|
|
const BASEBLOCKEX& b = blocks[idx];
|
|
if (src_pc < b.startpc || src_pc >= b.startpc + b.size * 4)
|
|
return false;
|
|
const uptr lo = b.fnptr;
|
|
const uptr hi = b.fnptr + b.x86size;
|
|
const auto range = links.equal_range(dst_pc);
|
|
for (auto it = range.first; it != range.second; ++it)
|
|
{
|
|
const uptr site = it->second.site & ~kLinkSiteCallBit;
|
|
if (site >= lo && site < hi)
|
|
return true;
|
|
}
|
|
return false;
|
|
}
|
|
|
|
// Test-only: number of registered link sites targeting dst_pc, and the
|
|
// total across all destinations. Used to pin that the map stays bounded
|
|
// by *live* sites rather than growing with every recompile.
|
|
size_t LinkCount(u32 dst_pc) const
|
|
{
|
|
const auto range = links.equal_range(dst_pc);
|
|
return static_cast<size_t>(std::distance(range.first, range.second));
|
|
}
|
|
|
|
size_t TotalLinkCount() const { return links.size(); }
|
|
#endif
|
|
};
|