felix86 26.10
This month we got a bunch of games working, and landed some pretty neat optimizations. Titles include Battle.net, EA, and even Denuvo games!
Gaming
Here’s a freshly captured video of some games on this new version!
Return stack buffer optimization
Our address cache keeps indirect translations really fast, and it has existed for a while.
However, there are two interesting observations:
- All indirect translations go through the shared address cache, putting a lot of pressure on the branch predictor
- Hardware almost always has a separate prediction mechanism for returns, usually called the return stack buffer
- This is because most functions in software will return to where they were called from
Initially, we thought of implementing these optimizations using the host stack, similar to Dolphin. However, this has some issues in our case:
- We want the host stack to always point to a
felix86_frameso that ourC.LDSP+C.JRinvalidation instructions workC.LDcan’t usegpwhich we allocate toThreadStatefor ptrace related reasons, leavingspas the only possible register- The invalidation must replace only one host instruction as it does now, otherwise rare races can occur
- We don’t use a
sigaltstack, so we can’t afford using signal handlers for overflow/underflow handling - Adding another source of potential emulator-side signals is not great
- Extra code per
CALL
The first problem can be solved by loading and storing a stack pointer relative to ThreadState. That adds even more instructions to CALL and RET. The second problem can be solved by using a ring buffer. However that too would add overhead on CALL and RET. Instead, after lots of consideration and measurements, we landed upon a different solution.
By inlining the address cache translation in RET, using JALR for CALL and RET for RET, we address both the aforementioned observations. We relieve a lot of branch predictor pressure from the address cache that comes from RET, and we use the hardware return stack buffer for predictions. This requires no extra stack, so no stack bookkeeping instructions in CALL/RET, and no extra signal handlers.
Importantly, we also don’t write to the address cache on CALL at runtime. Adding the instructions to write to the address cache on every CALL adds noticeable overhead. This would bring one to think that blocks should not end on CALL instructions, and x86 RET should do a RISC-V RET to the instructions after the CALL. This was our original plan, however it turns out some games wanted to make our lives harder.
For example, Celeste would have blocks that have garbage after the CALL. If we keep compiling that garbage, sometimes we would hit assertions in felix86 code, and sometimes we would even hit unmapped pages. So code after CALL can’t be considered compile-able for every program, even if it is for most.
Instead, we decided to add a linkable stub after each CALL that will jump to the next block. This way the RET instruction predicts correctly that it will return right after the host JALR, instead of jumping immediately to the next block and missing the hardware prediction. When the block is compiled, it will point to the stub after the CALL instead of to the block itself. This way, even if the block gets evicted from the address cache, it will continue to point to the stub and get re-inserted to the address cache and the RET will keep predicting correctly.
This optimization improves performance in many games, and is on by default. In TEKKEN 7, it seems to improve performance by about 10%.
EA Games
After recent fixes, the emulator is now able to run some EA games!
Need for Speed Heat can now run on felix86
Another EA game that now runs fine
Running these launchers makes you wish for a native Linux launcher. Or a native RISC-V one, but let’s not dream too big!
Denuvo
While trying to run NFS: Unbound, an EA game, it would hit an illegal instruction. We then noticed that this is a Denuvo game.
We knew what was wrong with emulation of Denuvo games for a while due to research done by FEX-Emu, but didn’t get around to implementing inline SMC, until now. Despite our implementation, NFS: Unbound continues to not work for reasons unknown, but there’s some Denuvo games that now work!
At this angle, you can see the monkey. Move the mouse, and you can’t
There’s some graphics bugs in this game, and I am not certain that it is a felix86 bug or a driver bug. But the fact that it gets in game means that our inline SMC implementation works, and other Denuvo games could work as well.
Now works, but performance isn’t great on this hardware
Debugging Denuvo games is hard though, as they may detect your attempts and lock you out. Here’s Persona 5 Royal, a game that doesn’t yet work, locking us out of debugging it to see why it crashes.
Sorry, no more debugging for today
Just because some Denuvo games work doesn’t mean that all of them do. So there’s still work to do.
Battle.net
After even more fixes, the Battle.net launcher can now run and you can play some Blizzard games! The launcher is somewhat finicky still, possibly related to unaligned atomics crossing the 8-byte boundary, which current RISC-V hardware can’t perform atomically. It may freeze during game download but the installation completes in the background and the games can run fine.
Diablo III on RISC-V with felix86
Some Hearthstone too!
Sekiro: Shadows Die Twice
There is an interesting discovery when trying to get this game to work. While the game would run fine with TSO disabled, it seemed to crash at random points while playing. The reason for this crash isn’t yet found. This happens when TSO is disabled, so enabling TSO was attempted. However, the game would then refuse to launch with TSO enabled.
The game works fine! Until it doesn’t…
This game uses Steam DRM, which runs some sort of decryption algorithm on the game’s code at runtime. This decryption algorithm must take less than 10 seconds, or an error window will show with the failure code S:0000065434, my guess being that S here stands for Slow. Without TSO emulation, felix86 is fast enough to finish the decryption in time. However, with TSO enabled, it times out and refuses to launch.
If more than 10 seconds passed during decryption, exit with an error
In most games, this doesn’t trigger. This part of the DRM code seems to be decrypting the .text region of the executable. For Sekiro, it is quite large at 43 MB, so with TSO enabled, and current RISC-V hardware, it is slow enough to fail. It seems that this behavior of Steam DRM is long documented, as for example in this blog post about steamdrmp.dll.
The solution to this is faster TSO emulation, which should come with future RISC-V hardware. In current hardware, TSO emulation currently slows down software by a varying amount, from 2x to 6x or more, depending on how much memory accesses play a role. So Sekiro debugging is paused for now.
Automatic latency and throughput measurements
We also implemented a latency and throughput measurer, which will measure latency and throughput for x86 instruction translations on real RISC-V hardware. This way we can sort and find the worst performing instructions, and also have data for how much our instruction translations improve over the months. This measurement code runs on our CI and makes new commits to update the measurements if there’s meaningful change.
MSVC strstr improvements
Monster Hunter World would take 10 minutes to load to menu. During its load time, the vast majority was spent on PCMPISTRI emulation, inside the MSVC strstr function over a big buffer. This is a rare instruction, one of the few instructions that we would use a function for, so it was very slow.
We now properly recompile this instruction using RVV. Compared to our slow function, this new version is 18x faster in the general case. Some of these gains translated to faster load times for Monster Hunter world, which can now load in 1 minute and 48 seconds, meaning it loads up 5.6x faster.
Fault on non-executable pages
In previous versions, felix86 would compile code as long as it was readable. Most games didn’t care, as they don’t jump to non-executable code – a fault would happen if they did. Sometimes, however, the fault is intentional.
Proton Wine will normally patch functions in steamclient.dll to direct those calls to a Linux library. Seven different app IDs are marked in Proton Wine’s source code to use a different mechanism: Instead of patching, mark the code as non-executable and change the RIP on segmentation fault. One of them is Call of Duty: Black Ops II, and the Zombies version as well. This is presumably done because those games will checksum steamclient.dll and refuse to work if it is tampered with. When FELIX86_GUEST_MEMORY_TRACKING_ENABLED is set, executing non-executable memory will now properly fault, and the game works.
Don’t forget to enable reduced_precision for this one
Executable profiles
Similar to last month’s Steam profiles, you can now set specific profiles per executable filename. The installer script will install a default profile to enable TSO for the EA launcher and other applications in ~/.config/felix86/profiles/executables. This feature can be useful to enable features for specific games that aren’t Steam games, or for specific executables within a Steam game, such as enabling TSO for just the EA launcher and not the whole game.
Here’s an example profile, from the documentation:
["TekkenGame-Win64-Shipping.exe"]
General.enabled_thunks = "glx,vk,wl"
Performance.unsafe_flags = true
Performance.inaccurate_minmax = true
SHA extension
We now emulate the SHA-1 and SHA-256 x86 instructions. RISC-V has no SHA-1 hardware acceleration extensions, so most of the x86 SHA-1 instructions are emulated with vector instructions, with the exception of SHA1RNDS4 which calls out to a function.
As for the SHA-256 instructions, SHA256RNDS2 is just a single VSHA2CL. However, the SHA256MSG1 and SHA256MSG2 instructions combined only do part of the work that VSHA2MS does. The equivalent would basically be a SHA256MSG1 + MOVDQA + PALIGNR + PADDD + SHA256MSG2. However, we don’t look for that sequence of instructions to replace with a single VSHA2MS, as it’s not how it is usually laid out in x86 software, such as in OpenSSL: The message schedule is interleaved with the rounds, and the SHA256MSG1 for a block runs many rounds before its SHA256MSG2. Still, VSHA2MS can help us emulate these instructions faster than the RVV-only alternative. For example, if we zero W[0] to W[4], place the destination of SHA256MSG2 in W[9] to W[12], and ensure W[14] and W[15] from the source are positioned correctly, we can use VSHA2MS to emulate the SHA256MSG2 instruction. In current hardware, this is measurably faster than the alternative, which would have a bad dependency chain, as W[18] and W[19] need the σ1(W[16]) and σ1(W[17]). For SHA256MSG1 however, zeroing isn’t enough, as the upper two results would get σ1(W[16]) and σ1(W[17]) added to them. Still, we can use the same instruction twice and use the lower two elements of the result as such:
vslidedown.vi temp, dst, 2 ; t = {W2, W3, -, -}
vslideup.vi temp, src, 2 ; t = {W2, W3, W4, W5}
vsha2ms.vv dst, vzero, vzero ; W0+σ0(W1), W1+σ0(W2)
vsha2ms.vv temp, vzero, vzero ; W2+σ0(W3), W3+σ0(W4)
vslideup.vi dst, temp, 2
These instructions should help accelerate programs that use checksuming, such as launchers that install games and verify file integrity afterwards.
Thanks for reading this post.
If you like this project, please give us a star on GitHub: https://github.com/OFFTKP/felix86
