Fix: DXSTG translates its tag instead of aiming a write-back at a host address

PCSX2's D-cache tag holds a host pointer -- CacheLine::load stores the
translated `ppf` -- where hardware holds a guest physical address, and
DXSTG copied CP0 TagLo into that field unmasked, so writeBackIfNeeded
dereferenced a word the guest chose. It now translates the tag page the
way a fill does and takes isValidPFN from the same translation; the
comment at the case has the rest.

That widens the write-back precondition: a DXSTG naming a resolvable
page arms a line that was never filled, where before only a fill could
set bit 11. That is the hardware behaviour.

The exposure was narrow. The wild store needed isValidPFN standing from
a real fill, and fills are gated by EnableEECache, which ships off; the
tag write itself was never gated. Cache.cpp is shared, so x86 was
affected identically, and both recCACHE() bodies are empty, so the
recompilers reach this through the interpreter.

DISABLED_DxstgDirtyStaysInsideGuestMemory graduates.
Pcsx2DxstgRedirectsAHostWrite, which pinned the defect, becomes
DxstgWriteBackTargetsTheTaggedGuestPage, and
DxstgOnAnUnresolvablePageDeclinesTheWriteBack is new.
Pcsx2AcceptsDirtyFromDxstgButCannotReportIt used the redirect as its
instrument and now reads guest memory.
This commit is contained in:
pstef
2026-08-08 16:02:13 +02:00
parent 75d78d8d5f
commit 7cae16e8e6
11 changed files with 225 additions and 290 deletions
+21 -1
View File
@@ -468,7 +468,27 @@ namespace R5900
const int way = addr & 0x1;
CacheLine line = cache.lineAt(index, way);
line.tag.setAddr(cpuRegs.CP0.n.TagLo);
// TagLo carries a guest physical page. Our tags do not: they hold the
// host pointer the fill translated to (CacheLine::load stores `ppf`),
// which is what writeBackIfNeeded dereferences, so copying the guest
// word in raw aimed a 64-byte store at an address the guest chose --
// setAddr zeroes the top 32 bits, so somewhere below 4 GiB: an
// emulator crash normally, or a write into whatever happened to be
// mapped there. Translate it the way a fill does instead, through the
// KSEG0 alias of the physical page (Memory.cpp maps 0x80000000 onto
// physical 0), so the write-back lands at the physical address the
// guest named, and take isValidPFN from the same translation so the
// two cannot disagree. A tag that does not resolve to plain memory --
// an MMIO handler page, or a physical address that does not exist --
// is marked unbacked; the line still caches and reports its flags,
// and loses its data on eviction (see the comment on CacheTag).
const u32 pageTag = cpuRegs.CP0.n.TagLo & ~static_cast<u32>(CacheTag::ALL_BITS);
const u32 alias = 0x80000000u | (pageTag & 0x1FFFFFFFu);
const VTLBVirtual vmv = vtlbdata.vmap[alias >> VTLB_PAGE_BITS];
const bool backed = !vmv.isHandler(alias);
line.tag.setValidPFN(backed);
line.tag.setAddr(backed ? vmv.assumePtr(alias) : static_cast<uptr>(pageTag));
line.tag.rawValue &= ~CacheTag::ALL_FLAGS;
line.tag.rawValue |= (cpuRegs.CP0.n.TagLo & CacheTag::ALL_FLAGS);
+14 -23
View File
@@ -470,9 +470,9 @@ void SQRT_S() {
// Invalid-operation keys off the SIGN BIT ALONE. -0 and the negative
// denormals raise it too, even though they are flushed to -0 and produce a
// perfectly ordinary +0: the exponent field plays no part. This used to sit
// inside the negative-normal arm below, so those two operand classes came
// back with FCR31 untouched. x86's recSQRT_S_xmm has always tested the sign
// perfectly ordinary +0: the exponent field plays no part. It used to sit
// inside a negative-normal arm, so those two operand classes came back with
// FCR31 untouched. x86's recSQRT_S_xmm has always tested the sign
// bit alone (iFPU.cpp, MOVMSKPS & 1), as has the FULL-mode DOUBLE path in
// iFPUd-arm64.cpp. Scored against a first-party capture over the sign x
// exponent matrix -- see EeRecFpu.SqrtSInvalidFlagFollowsTheSignBitAlone.
@@ -488,27 +488,18 @@ void SQRT_S() {
}
else if ( ( _FtValUl_ & 0x7F800000 ) == 0x7F800000 )
{
// Exponent 255 is an ORDINARY binade on the EE -- no Inf, no NaN, and
// the representable max is 0x7FFFFFFF, not FLT_MAX. So fpuDouble()'s
// clamp is not a rounding of this operand, it is a different operand,
// and the answer lands two binades low: sqrt(2^128) came back as
// 0x5F7FFFFF where the console gives 0x5F800000, and sqrt(+EEMAX) as
// 0x5F7FFFFF against 0x5FB504F3.
// Exponent 255 is an ordinary binade on the EE -- the representable max
// is 0x7FFFFFFF, not FLT_MAX -- so fpuDouble()'s clamp hands sqrt a
// different operand rather than a rounded one: sqrt(0x7FFFFFFF) came
// back 0x5F7FFFFF where the console gives 0x5FB504F3. Square-root
// |Ft|/4 and double it instead. 4 is an even power of two, so its own
// square root is exact and the sqrt below stays the only rounding step.
// recSQRT_S_xmm (iFPU-arm64.cpp) emits the same two steps and carries
// the rest of the argument.
//
// Square-root |Ft|/4 and double it. sqrt halves exponents, so the
// scaled operand (exponent field 253) and the doubled result are both
// ordinary representable singles -- no wider format is needed. 4 is an
// even power of two, so its own square root is exact and the identity
// contributes no rounding: the sqrt below is the only rounding step,
// exactly as on the untouched path. Same power-of-two prescale that
// ToDouble() uses to carry these operands into FULL mode
// (iFPUd-arm64.cpp), with the factor picked to suit sqrt so it can stay
// in single precision. The arm64 fast path emits the same two steps --
// see recSQRT_S_xmm in iFPU-arm64.cpp.
//
// RSQRT_S deliberately does NOT get this. Its two clamped operands
// currently cancel on rsqrt(2^128, 2^128); unclamping only the sqrt
// breaks that row. It is all-or-nothing and is a separate change.
// RSQRT_S does not get this: its two clamped operands cancel on
// rsqrt(2^128, 2^128), so unclamping only the sqrt breaks that row.
// It is all-or-nothing and is a separate change.
FPRreg quarter;
quarter.UL = ( _FtValUl_ & 0x7FFFFFFF ) - 0x01000000; // |Ft| / 4
_FdValf_ = 2.0 * sqrt( (double)quarter.f );
+4 -7
View File
@@ -1027,13 +1027,10 @@ static void recSQRT_S_xmm(int info)
// PS2 SQRT.S flag handling (interp SQRT_S, FPU.cpp; CHECK_FPU_EXTRA_FLAGS
// is always on): clear I|D unconditionally, then set I|SI whenever Ft's
// SIGN BIT is set. The exponent field plays no part — 0 and the negative
// denormals raise I|SI too, even though they flush to 0 and produce +0.
// This used to carry an extra `exp != 0` gate, which cost exactly those two
// operand classes their flag; x86's recSQRT_S_xmm (iFPU.cpp, MOVMSKPS & 1)
// and the FULL-mode DOUBLE path (iFPUd-arm64.cpp) have always tested the
// sign alone. Scored against a first-party capture over the sign × exponent
// matrix — see EeRecFpu.SqrtSInvalidFlagFollowsTheSignBitAlone.
// sign bit is set. The exponent field plays no part — -0 and the negative
// denormals raise I|SI too. x86's recSQRT_S_xmm tests MOVMSKPS & 1 the same
// way (iFPU.cpp), as does the FULL-mode DOUBLE path (iFPUd-arm64.cpp). See
// EeRecFpu.SqrtSInvalidFlagFollowsTheSignBitAlone.
// Read the Ft bits before Fabs clobbers EEREC_D, which may alias EEREC_T.
// GE-12: flag RMW on the resident FCR31; alloc first (eviction stores
// must precede the RWARG1 clobber and the branch arms). GE-20 gave SQRT
@@ -41,32 +41,24 @@
// `tag page | index << 6`, and a hit op through a second cached mapping of the
// same physical line hits.
//
// What it gets wrong is one defect, and it is worse than a wrong number.
// Round 2's finding was one defect, and it was worse than a wrong number.
// PCSX2's tag holds a *host* pointer where hardware holds a guest physical
// address, and DXSTG copies a guest-supplied 32-bit word straight into it
// (Cache.cpp, the DXSTG case). writeBackIfNeeded then dereferences that word.
// Pcsx2DxstgRedirectsAHostWrite maps a page at the address the guest asked
// for and watches 64 bytes land in it. Unmapped -- the normal case, since
// setAddr zeroes the top 32 bits so the target is always below 4 GiB -- it is
// an emulator crash instead.
// address, and DXSTG copied a guest-supplied 32-bit word straight into it, so
// writeBackIfNeeded dereferenced that word. It now translates the tag page and
// takes isValidPFN from the same translation (Cache.cpp, the DXSTG case); the
// write-back precondition that moves with it is at
// DxstgWriteBackTargetsTheTaggedGuestPage.
//
// What bounds it is EnableEECache, which PCSX2 ships off -- but indirectly,
// and the distinction matters when reading the tests. CHECK_CACHE gates only
// vtlb.cpp's guest load and store paths; the CACHE opcode itself and the cache
// entry points run whatever it is set to. The redirect needs isValidPFN (tag
// bit 11) set when the DXSTG lands, DXSTG's own mask (ALL_FLAGS = 0x7FF)
// cannot reach that bit and clear() drops it, so only a real fill can set it
// -- and a fill is what EnableEECache gates. With it off a guest's cache op
// still writes the host-pointer field and writeBackIfNeeded still declines,
// which is the precondition the never-filled and invalidated blocks of that
// test pin.
// EnableEECache, which PCSX2 ships off, bounded the old defect but only
// indirectly, and the distinction still matters when reading the tests.
// CHECK_CACHE gates only vtlb.cpp's guest load and store paths; the CACHE
// opcode itself and the cache entry points run whatever it is set to. The
// tests below therefore do not read the setting. They drive the model through
// its own entry points -- readCache32/writeCache32 and the interpreter's
// CACHE -- so everything here reproduces with EnableEECache at its default
// false.
//
// The tests below therefore do NOT read the setting. They drive the model
// through its own entry points -- readCache32/writeCache32 and the
// interpreter's CACHE -- so everything here, this defect included, reproduces
// with EnableEECache at its default false.
//
// Nothing here is fixed. Divergences are recorded from the real run.
// Everything else here is unfixed. Divergences are recorded from the real run.
#include <gtest/gtest.h>
@@ -288,21 +280,21 @@ TEST(EeCache2Console, DxstgSetsTheDirtyBitOnConsole)
//
// So the bit is measured by its consequence instead. The write-back below
// happens if and only if D survived the DXSTG, and it is the same sequence
// Pcsx2DxstgRedirectsAHostWrite uses -- with no DXLTG anywhere between the
// store and the eviction.
// DxstgWriteBackTargetsTheTaggedGuestPage uses -- with no DXLTG anywhere
// between the store and the eviction.
TEST(EeCache2Console, Pcsx2AcceptsDirtyFromDxstgButCannotReportIt)
{
u32 page = 0;
void* p = MapAt(&page);
if (!p)
GTEST_SKIP() << "could not map a page at any candidate host address";
// The consequence is observed in guest memory, at the physical page the tag
// names, which is where DXSTG now steers the write-back.
constexpr u32 kTargetPage = 0x00129000;
constexpr u32 kTarget = kTargetPage + kSetIndex * 64;
// Asking destroys it.
{
EeRecTestHarness h;
resetCache();
RunCacheOp(0x16, kProbeLine); // DXIN way 0, so no write-back can fire
cpuRegs.CP0.n.TagLo = page | kFlagDirty | kFlagValid;
cpuRegs.CP0.n.TagLo = kTargetPage | kFlagDirty | kFlagValid;
RunCacheOp(0x12, kProbeLine); // DXSTG
EXPECT_EQ(ReadTag(kProbeLine) & kFlagDirty, 0u)
<< "PCSX2's DXLTG now reports D; the console's 0x" << std::hex
@@ -310,32 +302,29 @@ TEST(EeCache2Console, Pcsx2AcceptsDirtyFromDxstgButCannotReportIt)
}
// Not asking shows it was there: the eviction only writes if D is set.
std::memset(p, 0xEE, 0x1000);
const u32* target = reinterpret_cast<const u32*>(
static_cast<uptr>(page) + kSetIndex * 64);
{
EeRecTestHarness h;
resetCache();
memWrite32(kTarget, 0xEEEEEEEEu);
writeCache32(kProbeLine, 0x5A5A0009u);
cpuRegs.CP0.n.TagLo = page | kFlagDirty | kFlagValid;
cpuRegs.CP0.n.TagLo = kTargetPage | kFlagDirty | kFlagValid;
RunCacheOp(0x12, kProbeLine); // DXSTG, D set
RunCacheOp(0x14, kProbeLine); // DXWBIN
EXPECT_EQ(target[0], 0x5A5A0009u) << "D did not survive the DXSTG";
EXPECT_EQ(memRead32(kTarget), 0x5A5A0009u) << "D did not survive the DXSTG";
}
// The control: the same sequence with D clear moves nothing, so the write
// above is attributable to the bit and not to DXWBIN writing regardless.
std::memset(p, 0xEE, 0x1000);
{
EeRecTestHarness h;
resetCache();
memWrite32(kTarget, 0xEEEEEEEEu);
writeCache32(kProbeLine, 0x5A5A0009u);
cpuRegs.CP0.n.TagLo = page | kFlagValid; // no D
cpuRegs.CP0.n.TagLo = kTargetPage | kFlagValid; // no D
RunCacheOp(0x12, kProbeLine);
RunCacheOp(0x14, kProbeLine);
EXPECT_EQ(target[0], 0xEEEEEEEEu) << "DXWBIN wrote back a clean line";
EXPECT_EQ(memRead32(kTarget), 0xEEEEEEEEu) << "DXWBIN wrote back a clean line";
}
munmap(p, 0x1000);
}
// The write-back target is the tag's page with the *set index* supplying the
@@ -434,77 +423,98 @@ TEST(EeCache2Console, Pcsx2HasNoInstructionCacheGeometry)
}
// ---------------------------------------------------------------------------
// DEFECT. Hardware's tag holds a guest physical address, so a write-back
// steered by DXSTG can only ever reach guest memory. PCSX2's tag holds a host
// pointer (CacheLine::load stores `ppf`), and DXSTG copies a guest word into it
// unmasked, so writeBackIfNeeded dereferences a guest-chosen host address.
// Hardware's tag holds a guest physical address, so a write-back steered by
// DXSTG can only ever reach guest memory. PCSX2's now does too (Cache.cpp, the
// DXSTG case).
//
// The precondition is exactly "the line is resident when the DXSTG lands":
// isValidPFN is bit 11 of the same word, DXSTG's mask (ALL_FLAGS = 0x7FF) does
// not reach it, and clear() drops it. Both halves are checked below, because
// "a guest can do this at will" and "a guest must first touch the line" are
// different findings and only one of them is true.
TEST(EeCache2Console, Pcsx2DxstgRedirectsAHostWrite)
// The old write-back precondition was "the line is resident when the DXSTG
// lands": isValidPFN is bit 11 of the same word, DXSTG's mask (ALL_FLAGS =
// 0x7FF) could not reach it, and clear() dropped it. isValidPFN now moves with
// the address, so a DXSTG naming a resolvable page arms the line whether or not
// it was ever filled. All three histories are exercised below.
TEST(EeCache2Console, DxstgWriteBackTargetsTheTaggedGuestPage)
{
// A physical page in main RAM, chosen to equal one of MapAt's candidates so
// the host page of the same number can be held open as a negative control.
constexpr u32 kTargetPage = 0x00129000;
constexpr u32 kTarget = kTargetPage + kSetIndex * 64;
u32 page = 0;
void* p = MapAt(&page);
if (!p)
GTEST_SKIP() << "could not map a page at any candidate host address";
ASSERT_EQ(page, kTargetPage) << "the control page moved; re-point kTargetPage";
std::memset(p, 0xEE, 0x1000);
const u32* target = reinterpret_cast<const u32*>(
const u32* host = reinterpret_cast<const u32*>(
static_cast<uptr>(page) + kSetIndex * 64);
{
EeRecTestHarness h;
resetCache();
memWrite32(kTarget, 0u);
memWrite32(kTarget + 4, 0u);
writeCache32(kProbeLine, 0x5A5A0009u);
writeCache32(kProbeLine + 4, 0xDEADBEEFu);
cpuRegs.CP0.n.TagLo = page | kFlagDirty | kFlagValid;
cpuRegs.CP0.n.TagLo = kTargetPage | kFlagDirty | kFlagValid;
RunCacheOp(0x12, kProbeLine); // DXSTG
RunCacheOp(0x14, kProbeLine); // DXWBIN
// 64 bytes of guest cache line, written to a host address the guest
// picked. On hardware this is a store to guest physical memory.
EXPECT_EQ(target[0], 0x5A5A0009u)
<< "PCSX2 no longer redirects the write; record the new behaviour";
EXPECT_EQ(target[1], 0xDEADBEEFu);
// 64 bytes of guest cache line, at the guest physical page the tag names.
EXPECT_EQ(memRead32(kTarget), 0x5A5A0009u);
EXPECT_EQ(memRead32(kTarget + 4), 0xDEADBEEFu);
EXPECT_EQ(host[0], 0xEEEEEEEEu) << "the write-back still reaches a host address";
}
// Never filled: bit 11 is clear and writeBackIfNeeded declines.
std::memset(p, 0xEE, 0x1000);
// Never filled, and filled-then-invalidated. Both used to be declined
// because bit 11 was clear; both now write back, carrying the cleared line.
for (const bool invalidate_first : {false, true})
{
SCOPED_TRACE(invalidate_first ? "filled then invalidated" : "never filled");
std::memset(p, 0xEE, 0x1000);
EeRecTestHarness h;
resetCache();
RunCacheOp(0x16, kProbeLine); // DXIN
cpuRegs.CP0.n.TagLo = page | kFlagDirty | kFlagValid;
memWrite32(kTarget, 0xA5A5A5A5u);
if (invalidate_first)
{
writeCache32(kProbeLine, 0xABCD1234u);
RunCacheOp(0x14, kProbeLine); // DXWBIN: writes back, then clears
}
else
{
RunCacheOp(0x16, kProbeLine); // DXIN
}
memWrite32(kTarget, 0xA5A5A5A5u);
cpuRegs.CP0.n.TagLo = kTargetPage | kFlagDirty | kFlagValid;
RunCacheOp(0x12, kProbeLine);
RunCacheOp(0x14, kProbeLine);
EXPECT_EQ(target[0], 0xEEEEEEEEu) << "an unfilled line also redirects";
}
// Filled and then invalidated: clear() drops bit 11 too.
std::memset(p, 0xEE, 0x1000);
{
EeRecTestHarness h;
resetCache();
writeCache32(kProbeLine, 0xABCD1234u);
RunCacheOp(0x14, kProbeLine); // DXWBIN: writes back, then clears
cpuRegs.CP0.n.TagLo = page | kFlagDirty | kFlagValid;
RunCacheOp(0x12, kProbeLine);
RunCacheOp(0x14, kProbeLine);
EXPECT_EQ(target[0], 0xEEEEEEEEu) << "an invalidated line also redirects";
EXPECT_EQ(memRead32(kTarget), 0u) << "the cleared line did not write back";
EXPECT_EQ(host[0], 0xEEEEEEEEu) << "the write-back still reaches a host address";
}
munmap(p, 0x1000);
}
// ---------------------------------------------------------------------------
// Tripwires. Each fails today and turns green when the defect above is fixed
// or the missing model appears.
// A DXSTG naming a page that does not resolve to plain guest memory leaves the
// line unbacked, so the write-back declines rather than dereferencing anything.
// 0x1FC00000-and-up is BIOS/unmapped territory at the top of the physical map.
TEST(EeCache2Console, DxstgOnAnUnresolvablePageDeclinesTheWriteBack)
{
EeRecTestHarness h;
resetCache();
writeCache32(kProbeLine, 0x5A5A0009u);
cpuRegs.CP0.n.TagLo = 0x1FFFF000u | kFlagDirty | kFlagValid;
RunCacheOp(0x12, kProbeLine); // DXSTG
RunCacheOp(0x14, kProbeLine); // DXWBIN -- must be a no-op, not a store
// Reaching here without a fault is the assertion; the flags still round-trip.
EXPECT_EQ(ReadTag(kProbeLine) & (kFlagValid | kFlagDirty), 0u);
}
TEST(EeCache2Console, DISABLED_DxstgDirtyStaysInsideGuestMemory)
// ---------------------------------------------------------------------------
// Tripwires. DxstgDirtyStaysInsideGuestMemory has graduated and holds; the rest
// fail today and turn green when the missing model appears.
TEST(EeCache2Console, DxstgDirtyStaysInsideGuestMemory)
{
u32 page = 0;
void* p = MapAt(&page);
@@ -546,7 +556,8 @@ TEST(EeCache2Console, DISABLED_InstructionCacheGeometryIsModelled)
TEST(EeCache2Console, DISABLED_AllEeCache2MatchesConsole)
{
// Graduation: everything above that is currently recorded as a divergence
// has to hold at once.
// has to hold at once. The DXSTG half already does; the instruction cache
// is what keeps this disabled.
u32 page = 0;
void* p = MapAt(&page);
if (!p)
@@ -489,11 +489,11 @@ TEST(EeCacheConsole, InvalidateKeepsThePhysicalTagOnConsole)
// is the observation, which is unchanged; the name says 0x78 because that is
// what the observation plus round 2 amounts to.
//
// Its own bit 11 (isValidPFN) is NOT reachable this way, which was worth
// checking rather than assuming: setAddr() masks with ALL_BITS = 0xFFF and so
// preserves bit 11 across the store, and flags() masks it back off on the way
// out, so a guest DXSTG can neither set nor read it. It can still be left
// standing from a fill, and round 2 shows what that costs.
// Its own bit 11 (isValidPFN) is not reachable as a *bit*: DXSTG's flag mask is
// ALL_FLAGS = 0x7FF, setAddr() masks with ALL_BITS = 0xFFF and so preserves bit
// 11 across the store, and flags() masks it back off on the way out, so a guest
// can neither read it nor write it directly. It is now set from whether the tag
// DXSTG stored resolves to guest memory; round 2 has why.
TEST(EeCacheConsole, DxstgWritableMaskIsPtagLoPlus0x78)
{
EeRecTestHarness h;
@@ -299,17 +299,9 @@ struct FlagSituation
u32 fcr31;
bool check_fd;
u32 fd;
// The arm64 recompiler now raises O/SO (fpuEmitOverflowFlags in
// iFPU-arm64.cpp), so "Overflow" matches silicon on both engines. "NAN
// math" does not, and cannot at this clamp mode: it feeds ADD.S two raw
// exp-255 words, which the interpreter turns into ±fMax through fpuDouble
// (giving an Inf sum, hence O) while the default fast path hands them to
// the host untouched and gets a NaN, which is not an overflow. Turning on
// CHECK_FPU_EXTRA_OVERFLOW makes the JIT clamp its operands the same way
// and the row aligns — measured by
// EeFpuFcrConsoleConformance.NanMathOverflowIsAnOperandClampModeDifference.
// So this flag now means "diverges at the default clamp mode", not "the
// recompiler has no O/SO".
// Rows where the fast path does not reproduce the console word. "NAN
// math" diverges only at the default clamp mode; see
// DISABLED_NanMathOverflowIsAnOperandClampModeDifference.
bool bad_jit;
};
constexpr FlagSituation kFlagSituations[] = {
@@ -443,7 +435,7 @@ TEST(EeFpuFcrConsoleConformance, DISABLED_EnginesAgreeExceptOnTheOverflowFlags)
<< "the engines now AGREE. If the JIT started clamping its operands "
"at the default clamp mode, drop this row from "
"kFcrEngineDivergences.";
// Pin WHICH side is right, so a future "fix" that aligns them by
// Pin which side is right, so a future "fix" that aligns them by
// removing the interpreter's flags fails here instead of passing.
EXPECT_EQ(got[0], s.fcr31)
<< "[interp] must stay the console-matching side";
@@ -509,8 +501,7 @@ TEST(EeFpuFcrConsoleConformance, DISABLED_NanMathOverflowIsAnOperandClampModeDif
"fpuDouble does must reproduce the interpreter's O|SO exactly -- if "
"this fails the divergence is NOT the operand-clamp axis and the "
"attribution above is wrong";
// The value is identical in all three legs, which is precisely why this
// divergence stayed invisible until FCR31 was read back.
// The value is identical in all three legs; only FCR31 moves.
EXPECT_EQ(res[0], 0x7F7FFFFFu);
EXPECT_EQ(res[1], res[0]);
EXPECT_EQ(res[2], res[0]);
@@ -519,36 +510,17 @@ TEST(EeFpuFcrConsoleConformance, DISABLED_NanMathOverflowIsAnOperandClampModeDif
// ---------------------------------------------------------------------------
// The O/U class, engine against engine.
//
// The two console rows above are one window into a whole family: FCR31's
// overflow and underflow maintenance, which pcsx2/FPU.cpp performs on EVERY
// arithmetic op and the recompilers performed on none. Three distinct
// behaviours live in the interpreter and all three are testable without a
// capture, because the interpreter is the reference side here:
// The console rows above are one window into a family: the FCR31 overflow and
// underflow maintenance pcsx2/FPU.cpp performs on every arithmetic op and the
// recompilers perform on none. It needs no capture, because the interpreter is
// the reference side; the three behaviours it implements are the three groups
// of kFamCases below.
//
// 1. the ten ops that call checkOverflow(result, O|SO) and then
// checkUnderflow(result, U|SU) -- ADD/SUB/MUL, the A-forms ADDA/SUBA/
// MULA, and the multiply-accumulates MADD/MSUB/MADDA/MSUBA. On an Inf
// result they set O|SO and RETURN, leaving U alone; otherwise they CLEAR
// O and then clear (or set) U.
// 2. ABS/NEG/MAX/MIN, which clearFPUFlags(O|U) and nothing else.
// 3. DIV/SQRT/RSQRT, which touch I and D but pass 0 to checkOverflow, so
// O and U must survive them untouched. These are the negative controls,
// and they are live ones: rows 1 and 2 in the same table prove the probe
// can see an FCR31 change at all, so "unchanged" here means preserved
// rather than unobserved.
//
// The clear in (1) and (2) is observable on its own -- preset O and U through
// ctc1 (both are in the writable mask) and run a non-overflowing op. That is
// why the pre-state below is 0x0100C001 rather than the bare fixed-ones word:
// it makes set, clear and preserve three distinguishable outcomes instead of
// two.
//
// The underflow half of (1) is deliberately NOT exercised here: FZ is set in
// every FP environment PCSX2 runs the EE under (DAZ+FTZ+ChopZero is the
// shipping default, and both ScopedFpEnv kinds this file uses keep FZ on), so
// no denormal result can reach checkUnderflow and U is only ever cleared. What
// happens with FZ off is the denormal-operand question -- a separate work
// item -- and is pinned by DISABLED_UnderflowFlagsNeedFzOff below.
// Only the overflow half is exercised. FZ is set in every FP environment PCSX2
// runs the EE under -- DAZ+FTZ+ChopZero is the shipping default and both
// ScopedFpEnv kinds this file uses keep FZ on -- so no denormal result can
// reach checkUnderflow and U is only ever cleared.
// DISABLED_UnderflowFlagsNeedFzOff pins the FZ-off half.
namespace
{
enum FamOp
@@ -581,7 +553,7 @@ struct FamCase
u32 want_fcr31;
};
// No row here overflows the intermediate PRODUCT of a multiply-accumulate:
// No row here overflows the intermediate product of a multiply-accumulate:
// that corner is a deliberate default-clamp-mode divergence between the
// engines (see recMADD_S_xmm in iFPU-arm64.cpp, pinned by
// EeRecFpu.MaddSProductOverflowDefaultModeMatchesX86Jit), and pulling it in
@@ -623,7 +595,9 @@ constexpr FamCase kFamCases[] = {
{"MIN.S clears O|U", FA_MIN, 0, kOne, kTwo, kFcr31FixedOnes},
// (3) Negative controls -- the divide unit passes 0 to checkOverflow, so
// O and U must come out exactly as they went in.
// O and U must come out exactly as they went in. Groups (1) and (2) above
// move the same bits from the same pre-state, so "unchanged" here means
// preserved rather than unobserved.
{"DIV.S preserves", FA_DIV, 0, kOne, kTwo, kOuPreset},
{"SQRT.S preserves", FA_SQRT, 0, 0, kFour, kOuPreset},
{"RSQRT.S preserves", FA_RSQRT, 0, kOne, kFour, kOuPreset},
@@ -700,14 +674,11 @@ TEST(EeFpuFcrConsoleConformance, DISABLED_EnginesAgreeOnOverflowFlagsAcrossTheAr
const u32 jit = RunFamCase(c, true, &res[1]);
SCOPED_TRACE(::testing::Message() << c.what);
// The interpreter is the reference side: it is what matches the console
// on the two captured rows above, so pin it to the value FPU.cpp's
// checkOverflow/clearFPUFlags model says it must produce.
// Pin the interpreter to what FPU.cpp's checkOverflow/clearFPUFlags
// model says it must produce.
EXPECT_EQ(interp, c.want_fcr31) << "[interp] no longer matches the "
"checkOverflow model in FPU.cpp";
EXPECT_EQ(jit, interp) << "engines disagree on FCR31 O/U";
// Only the flags are supposed to be under test -- if the arithmetic
// diverged too, the row is measuring the wrong thing.
EXPECT_EQ(res[1], res[0]) << "engines disagree on the RESULT, so this "
"row no longer isolates the flag write";
++checked;
@@ -715,13 +686,10 @@ TEST(EeFpuFcrConsoleConformance, DISABLED_EnginesAgreeOnOverflowFlagsAcrossTheAr
EXPECT_EQ(checked, kFamCaseCount);
}
// Several flag writers in ONE block, which is where the recompiler's FCR31
// block residency (GE-12) has to hold the whole model together: the arithmetic
// family now read-modify-writes the same allocator-resident FCR31 that C.cond
// writes the condition bit into, so a bad mask would either eat C or make SO
// non-sticky. Both orderings are checked because they exercise different
// halves: O has to come back down when a later op does not overflow, and it has
// to stay up when the last one does. Neither may disturb C.
// Several flag writers in one block, where the recompiler's FCR31 block
// residency (GE-12) has to hold the model together: the arithmetic family
// read-modify-writes the same allocator-resident FCR31 that C.cond writes the
// condition bit into, so a bad mask would either eat C or make SO non-sticky.
// TRIPWIRE -- see the O/SO rejection note above.
TEST(EeFpuFcrConsoleConformance, DISABLED_OverflowFlagsComposeAcrossOneBlock)
{
@@ -776,7 +744,7 @@ TEST(EeFpuFcrConsoleConformance, DISABLED_OverflowFlagsComposeAcrossOneBlock)
// The underflow half of checkUnderflow(result, U|SU), which needs a denormal
// result and therefore needs FZ off. DISABLED because it is the denormal-
// operand work item, not this one: with FZ off the two engines also disagree
// on the VALUE (the interpreter flushes the denormal to signed zero inside
// on the value (the interpreter flushes the denormal to signed zero inside
// checkUnderflow, the recompilers keep it), and pinning the flag without the
// value would assert half a behaviour. Force-enable to see the current state.
TEST(EeFpuFcrConsoleConformance, DISABLED_UnderflowFlagsNeedFzOff)
@@ -24,14 +24,10 @@
// engine-vs-engine divergence is not. The console column is therefore carried
// as data and asserted only by the DISABLED tripwire at the bottom.
//
// SQRT.S is the one op that has left this compromise. Exponent 255 is an
// ordinary binade, so its operands never needed saturating at all: both engines
// now compute sqrt(|Ft|/4)*2 and match the console exactly, without widening to
// double and without changing any operand whose exponent field is <= 254. That
// is what moved rows 44 and 45 out of the value-only column below. The same
// argument is available to ABS.S (see the DISABLED EeFpuAbsNegClamp tripwire)
// but NOT to the arithmetic ops, whose results genuinely exceed what the host
// single can hold.
// SQRT.S has left this compromise: both engines scale instead of clamping and
// match the console, which is what moved rows 44 and 45 out of the value-only
// column below. ABS.S likewise (ee_fpu_absneg_clamp_tests.cpp). The arithmetic
// ops cannot follow: their results do exceed what a host single can hold.
//
// The measured console divergences, all shared by both engines and all
// deliberate, for the record:
@@ -235,11 +231,9 @@ TEST(EeFpuOverflowConsole, EnginesAgreeExceptOnTheDocumentedRows)
}
// ---------------------------------------------------------------------------
// ENABLED. Classifies the divergence list by measurement rather than by
// assertion in a comment: every listed row must close when the operand clamp
// is on. The else-branch is the liveness clause for any future entry that does
// NOT close -- a defect rather than the mode axis, which is what rows 44/45
// were before SQRT gained its clamp.
// ENABLED. Every listed row must close when the operand clamp is on. The
// else-branch covers a future entry that does not close: that would be a
// defect rather than the mode axis.
// ---------------------------------------------------------------------------
TEST(EeFpuOverflowConsole, OperandClampHealsEveryDocumentedDivergence)
{
@@ -300,32 +294,20 @@ TEST(EeFpuOverflowConsole, DefaultClampModeSaturatesToFltMaxOnBothEngines)
}
// ---------------------------------------------------------------------------
// REGRESSION TEST for the defect this capture surfaced, and then for the fix.
// Both engines against the console on every SQRT row in the capture.
//
// Round one: recSQRT_S_xmm was the one emitter in iFPU-arm64.cpp that never
// clamped its operand, so an exponent-255 Ft reached Fsqrt as a host +Inf and
// fpuClampResult flattened the result to 0x7F7FFFFF, while the interpreter's
// sqrt(fpuDouble(Ft)) landed two binades away at 0x5F7FFFFF. SQRT was given a
// clamp to match, and the engines agreed -- on 0x5F7FFFFF, which is not what
// the console returns either. Agreement is a weaker property than accuracy and
// that round bought it at the cost of accuracy.
// The capture surfaced this as a clamp defect: recSQRT_S_xmm was the one
// emitter in iFPU-arm64.cpp that never clamped its operand, so an exponent-255
// Ft reached Fsqrt as a host +Inf and fpuClampResult flattened the result to
// 0x7F7FFFFF, while the interpreter's sqrt(fpuDouble(Ft)) gave 0x5F7FFFFF.
// Clamping SQRT too made both engines say 0x5F7FFFFF, which is not the console
// value either. They scale instead now -- see SQRT_S (pcsx2/FPU.cpp) and
// recSQRT_S_xmm (pcsx2/arm64/iFPU-arm64.cpp).
//
// Round two, what this now pins: neither engine clamps this operand. Both
// compute sqrt(|Ft|/4)*2, which keeps operand and result inside the ordinary
// single range without widening to double -- exponent 255 is an ordinary binade
// on the EE, so there was never anything here to saturate. See SQRT_S
// (pcsx2/FPU.cpp) and recSQRT_S_xmm (pcsx2/arm64/iFPU-arm64.cpp).
//
// Before the fix this failed on rows 44 and 45 (console 5fb504f3 / 5f800000,
// both engines 5f7fffff) and passed on row 46, whose Ft has exponent field 254
// and so never reached the clamp. Row 46 is therefore the negative control for
// the scaling branch's condition: if the branch were simply always taken, or
// the condition inverted, row 46 would move.
//
// The console value is asserted, not merely engine agreement -- agreement can
// always be reached by degrading whichever engine is nearer silicon, which is
// how round one went wrong. Both clamp modes are checked because the old clamp
// was gated on CHECK_FPU_OVERFLOW and the replacement deliberately is not.
// Rows 44 and 45 failed until then (console 5fb504f3 / 5f800000, both engines
// 5f7fffff); row 46's Ft has exponent field 254, never reached the clamp, and
// passed throughout. Both clamp modes are checked because the clamp this
// replaced was gated on CHECK_FPU_OVERFLOW.
// ---------------------------------------------------------------------------
TEST(EeFpuOverflowConsole, SqrtMatchesConsoleOnEveryCapturedOperand)
{
@@ -374,28 +356,19 @@ TEST(EeFpuOverflowConsole, SqrtMatchesConsoleOnEveryCapturedOperand)
}
// ---------------------------------------------------------------------------
// The same property as above, over the WHOLE exponent-255 class rather than the
// three patterns the capture happens to contain.
// The same property over the whole exponent-255 class rather than the three
// patterns the capture happens to contain. As host bit patterns those words
// are infinities, quiet NaNs and signalling NaNs; to the EE they are all large
// finite floats, so the pool carries every shape. The signalling ones are the
// half a host-NaN-aware implementation gets wrong -- see recSQRT_S_xmm
// (iFPU-arm64.cpp) for why the clamp they replaced had to be an integer Umin
// rather than an Fminnm.
//
// This exists because the class splits on an axis the capture cannot see. As
// HOST bit patterns, exponent-255 words are infinities, quiet NaNs and
// signalling NaNs; to the EE they are all just large finite floats. The old
// arm64 clamp had to be an integer Umin rather than an Fminnm precisely because
// of that split -- FMINNM only prefers the number against a QUIET NaN, while a
// signalling operand comes back merely quieted, so half the mantissa space
// (4194303 of the 8388608 positive patterns) would have passed through a clamp
// that was supposed to catch it. Testing the exponent FIELD, as both engines
// now do, never asks the host what kind of NaN it thinks it is holding, so the
// whole taxonomy should be irrelevant -- and this is what proves it.
//
// Expected values are correctly-rounded square roots computed by exact integer
// arithmetic (math.isqrt on the significand, round-to-nearest-even, which is
// the divide unit's mode), NOT by a host float, so they cannot inherit the
// behaviour under test. That model was validated against silicon on the six
// operands the capture does witness -- marked `true` below -- and agreed on all
// six including the exponent-254 control. The three unwitnessed rows are
// therefore computed expectations, not measurements; they are here for class
// coverage and are flagged as such.
// The wanted values are correctly-rounded square roots computed by exact
// integer arithmetic (math.isqrt on the significand, round-to-nearest-even,
// the divide unit's mode) rather than by a host float. The six marked `true`
// were read off silicon; the other three are computed only, for class
// coverage.
// ---------------------------------------------------------------------------
TEST(EeFpuOverflowConsole, SqrtMatchesConsoleOnEveryExponent255Operand)
{
@@ -421,8 +394,7 @@ TEST(EeFpuOverflowConsole, SqrtMatchesConsoleOnEveryExponent255Operand)
{0x7FC00000u, 0x5F9CC471u, 0x5F9CC470u, true, "exp255 mant 0x400000 (host +qNaN, smallest)"},
{0x7FFFFFFFu, 0x5FB504F3u, 0x5FB504F2u, true, "+EEMAX (host +qNaN, largest)"},
{0xFFFFFFFFu, 0x5FB504F3u, 0x5FB504F2u, true, "-EEMAX (host -qNaN, largest)"},
// CONTROL: exponent field 254, so the scaling branch must NOT fire.
// If it does, this row comes back one binade low.
// Control: exponent field 254, below the scaling branch.
{0xFF7FFFFFu, 0x5F7FFFFFu, 0x5F7FFFFFu, true, "-FLT_MAX (exp 254 -- CONTROL)"},
};
@@ -471,7 +443,7 @@ TEST(EeFpuOverflowConsole, SqrtMatchesConsoleOnEveryExponent255Operand)
// ---------------------------------------------------------------------------
// TRIPWIRE for the later hardware-alignment stage. Both engines, every row,
// against the console. Expected to fail on 37 of 57 rows today for the three
// against the console. Expected to fail on 35 of 57 rows today for the three
// documented reasons at the top of this file.
// ---------------------------------------------------------------------------
TEST(EeFpuOverflowConsole, DISABLED_AllRowsMatchConsole)
@@ -13,8 +13,8 @@
// CHECK_FPU_GUARDED / fpuGuardedAddSub option, which is ON by default (games
// like True Crime NYC and Jak 3 misrender without it, and per-game flagging
// proved impractical) but can be turned off globally for EE-heavy titles that
// don't need it. These tests run under the default (ON); DisableEmitsPlainOp
// at the bottom pins the opt-out path.
// don't need it. These tests run under the default (on);
// DisableEmitsPlainOpMatchingInterp at the bottom pins the opt-out path.
//
// THESE ARE JIT-ONLY TESTS. The shared interpreter's ADD_S/SUB_S (FPU.cpp) is a
// plain host float + float (fpuDouble() returns float and does no masking),
@@ -15,13 +15,10 @@
// - Negative nonzero divisor: interp rounds sqrt(|Ft|) into a float temp
// before dividing, so its divide is single-precision and matches native
// bit-for-bit. I|SI raised.
// - Positive nonzero divisor: interp computes Fs / sqrt(Ft) in DOUBLE (bare
// libm sqrt returns double, so the divide promotes) then rounds to float,
// while native/x86/hardware stay single-precision. The results disagree by
// exactly <=1 ULP on inexact quotients (~10% of random positive-Ft inputs).
// Native reproduces the single-precision EE FPU / x86 result, so those
// values are asserted JIT-only; the interp double-rounding divergence is
// pinned as a DISABLED tripwire (RsqrtSPositivePathDivergesFromInterp).
// - Positive nonzero divisor: same, since RSQRT_S() (pcsx2/FPU.cpp) stopped
// dividing by the double libm sqrt returns. At the production rounding
// mode the two still part by one ULP, pinned by
// DISABLED_RsqrtSPositivePathDivergesInProductionFpEnv.
#include "harness/EeRecTestHarness.h"
@@ -81,7 +78,7 @@ u32 fuzzOperand(Lcg& r)
// ---------------------------------------------------------------------------
// Differential fuzzer over the exactly-matching domain: zero and negative
// divisors (interp and native both stay single-precision there). Any Fs.
// Run()'s auto-diff checks the result value; the sticky flags are diffed too.
// The result value and the sticky flags are both diffed.
// ---------------------------------------------------------------------------
// These three ran green only because the old harness rounded to nearest. Under
// the production environment the interpreter's double-rounded RSQRT lands one
@@ -429,8 +429,8 @@ TEST(EeRecFpu, CvtWNegativeNanSaturatesToIntMin)
// ----- SQRT.S sticky-flag handling -----------------------------------
//
// PS2 SQRT.S clears the I|D cause flags unconditionally and sets I|SI when
// Ft is negative non-zero (interp SQRT_S, FPU.cpp; CHECK_FPU_EXTRA_FLAGS is
// PS2 SQRT.S clears the I|D cause flags unconditionally and sets I|SI whenever
// Ft's sign bit is set (interp SQRT_S, FPU.cpp; CHECK_FPU_EXTRA_FLAGS is
// hardcoded on). Run()'s auto-diff does not gate on fprc[31], so assert the
// flag bits directly on both snapshots (they must agree — the JIT matches
// interp). Result value (sqrt(|Ft|)) is unchanged and stays in the auto-diff.
@@ -478,13 +478,11 @@ TEST(EeRecFpu, SqrtSPositiveClearsStaleIDFlags)
// Found by a randomized SQRT.S differential, which is why a case this small
// went unnoticed: every hand-written SQRT.S test above uses +/-4.0.
//
// The FLAG half of this used to be asserted here as "no I|SI, the zero path is
// not the negative path", on nothing but the two engines agreeing. That was
// wrong: ps2autotests' sqrt.expected prints results only, never FCR31, so it
// could not have said either way. A first-party capture that does record FCR31
// (fpmatrix cases 226/227) shows the console raising I|SI here. It now lives
// in SqrtSInvalidFlagFollowsTheSignBitAlone below, which owns the whole rule;
// this test keeps the value and asserts only that D stays clear.
// The flag half used to be asserted here as "no I|SI", on nothing but the two
// engines agreeing; ps2autotests' sqrt.expected prints results only, never
// FCR31. The console does raise I|SI on -0, and the rule now lives in
// SqrtSInvalidFlagFollowsTheSignBitAlone below. This test keeps the value and
// asserts only that D stays clear.
TEST(EeRecFpu, SqrtSOfNegativeZeroIsPositiveZero)
{
EeRecTestHarness h;
@@ -500,16 +498,12 @@ TEST(EeRecFpu, SqrtSOfNegativeZeroIsPositiveZero)
}
// ----- SQRT.S's I flag keys off the sign bit, not the exponent --------
// The rule, from a first-party PS2 capture that records FCR31 alongside the
// result (~/.claude/projects/.../captures/fpmatrix, corpus 62dd6882, the 38
// SQRT.S cases):
//
// SQRT.S raises I|SI whenever Ft's SIGN BIT is set. Full stop. The
// exponent field plays no part.
//
// So -0.0 and the negative denormals -- which are flushed to -0 before the op
// and produce +0, a perfectly ordinary result -- still raise invalid-operation.
// Every capture row agrees; the ten below are the sign x exponent matrix:
// From a first-party PS2 capture that records FCR31 alongside the result
// (corpus 62dd6882, the 38 SQRT.S cases): SQRT.S raises I|SI whenever Ft's
// sign bit is set, whatever the
// exponent. So -0.0 and the negative denormals, flushed to -0 before the op
// and producing an ordinary +0, still raise invalid-operation. Every capture
// row agrees; the ten below are the sign x exponent matrix:
//
// case 226 sqrt 00000000 -> 00000000/01000001 +0
// case 227 sqrt 80000000 -> 00000000/01020041 -0 <- I|SI
@@ -523,23 +517,9 @@ TEST(EeRecFpu, SqrtSOfNegativeZeroIsPositiveZero)
// case 241 sqrt 80800000 -> 20000000/01020041 -MIN_NORMAL <- I|SI
//
// 0x01020041 is I|SI plus FCR31's two always-set bits (0x01000001); 0x01000001
// is those two bits alone. The positive rows are CONTROLS and they are why this
// is a matrix rather than two rows: the fix is a deletion (a gate on the
// exponent field comes out of the flag test), and a deletion that went too far
// -- dropping the sign test as well -- would raise I on every SQRT.S. Nothing
// but the positive rows would notice.
//
// Those controls are live, not decorative, and that was measured rather than
// assumed: with the sign test also deleted from both engines, all six positive
// rows fail on both and the four negative rows still pass. The two -0/-denormal
// rows were likewise seen failing before the fix (0x01000001 against
// 0x01020041, both engines) and passing after.
//
// Both engines used to gate the flag on `exp != 0 && sign`, so both missed
// exactly the two rows where the sign is set and the exponent is zero. x86's
// recSQRT_S_xmm (iFPU.cpp) has always tested MOVMSKPS's sign bit alone, as has
// the FULL-mode DOUBLE path in iFPUd-arm64.cpp -- upstream's x86 JIT is the
// only column in the capture that got these two rows right.
// is those two bits alone. The positive rows are controls: a deletion that took
// the sign test out along with the exponent gate would raise I on every SQRT.S,
// and nothing but those rows would catch it.
namespace {
struct SqrtFlagRow { u32 ft, result, fcr31; const char* what; };
constexpr SqrtFlagRow kSqrtFlagRows[] = {
@@ -912,13 +892,11 @@ TEST(EeRecFpu, MsubaSSubtractsProductFromAccumulator)
// ----- MADDA/MSUBA must NOT clamp the intermediate product -----------
//
// Interp MADD_S/MSUB_S route the fs*ft product through fpuDouble (clamping it
// to +-fMax) before the accumulate, but MADDA_S/MSUBA_S (FPU.cpp) add the raw
// product directly and overflow-check only the final ACC. Clamping the product
// in all four ops diverges when fs*ft overflows: an overflowing product
// clamped to +fMax cancels against an opposite-signed ACC (-> 0) instead of
// overflowing the accumulate (-> +-fMax). Run()'s auto-diff compares ACC, and
// these cases are chosen so JIT and interp agree only without the product clamp.
// A product clamped to +-fMax cancels against an opposite-signed ACC (-> 0)
// instead of overflowing the accumulate (-> +-fMax), and these cases are chosen
// so that only the unclamped behaviour survives. Neither engine clamps it any
// more; the interpreter rounds the product to an EE single and stops there if
// it overflowed (eeMulAccumulate).
// The next four all turn on the raw product reaching the accumulator as Inf,
// which only happens at round-to-nearest: round-toward-zero, the production
// mode, rounds an overflowing product to +/-FLT_MAX instead. Under that mode
@@ -277,7 +277,8 @@ TEST(EeRecMmi, PexcwAliasedRdEqualsRt)
TEST(EeRecMmi, QfsrvAdjacentSourceContiguous)
{
// Rs == Rt+1 (a1 == a0+1) hits the contiguous-memory path that reads the
// two source registers directly and skips the temp-buffer stores. sa = 4 bytes.
// two source registers directly instead of going through the general
// path's TBL. sa = 4 bytes.
EeRecTestHarness h;
h.SetMmiPair(reg::a0, 0x1122334455667788ull, 0x99AABBCCDDEEFF00ull); // Rt
h.SetMmiPair(reg::a1, 0xAABBCCDD11223344ull, 0x5566778899AABBCCull); // Rs
@@ -288,7 +289,7 @@ TEST(EeRecMmi, QfsrvAdjacentSourceContiguous)
TEST(EeRecMmi, QfsrvNonAdjacentSource)
{
// Rs != Rt+1 (a2 != a0+1) takes the temp-buffer path. Same Rt/Rs values,
// Rs != Rt+1 (a2 != a0+1) takes the general TBL path. Same Rt/Rs values,
// sa = 7 bytes.
EeRecTestHarness h;
h.SetMmiPair(reg::a0, 0x1122334455667788ull, 0x99AABBCCDDEEFF00ull); // Rt