Box86/Box64 vs. QEMU vs. FEX (Vs Rosetta2)
box86.org
box86.org
With box86/box64, I've been able to run Steam and even Wine/Proton with DXVK translating DirectX to Vulkan. I can even run older 3D Windows games like Skyrim! Though it did glitch on the infamous cart intro.
Do you have some clean distro that boots into something usable without mouse/keyboard? Some documentation for first stage boots? Gits?
https://github.com/ProjectValhalla/OdinMultiBootGuides
I don't think it's ready yet for full time use, as the joystick is mapped incorrectly for most games, but something to keep an eye on.
https://liliputing.com/2022/06/compare-handheld-gaming-pc-sp...
Realistically, an x87-specific JIT could do significant instruction reordering, lift/reoptimize the underlying code (much of existing x87 code was compiled a very long time ago on older compilers), and vectorize the underlying integer float emulation, or even trace and move some computation to another core or a coprocessor like a GPU or DSP (often idle in embedded cpus).
Many games work fine with x87 lowered to 64-bit or even 32-bit floats, and depending on the workload there's a middle ground where you could understand (or approximate) the current level of precision error for a value, generally run at a lower precision, and trace operations / "catch up" on precision at batched intervals.
You could also emulate it by arbitrarily dropping precision, but as a translator that means breaking bincompat, and more importantly breaking programs the use 80bit format (a lot of fortran).
Obviously many games (especially old ones) perform fine as they’re only using 80bit because at the time x87 was the only hardware fp available on x86 hardware, not because they needed that perf.
Even lowering the precision of the x87 unit isn’t sufficient as that only reduces the precision of the mantissa not the exponent.
Even outside of the core arithmetic (excluding negation which is really easy in all ieee754 formats) there is a whole bunch of state that you need to keep track of to ensure identical behavior.
Obviously if you are willing to break precision guarantees, etc then breaking state isn’t a problem, but if you’re trying to be something like Rosetta - eg completely general and running anything - you don’t really have the freedom to do that.
** sorry skim reading I missed your 128bit and x87 perf questions. Yes an emulator can (should?) use hw 128bit for the arithmetic if it’s available but on vast majority of hardware it isn’t.
You are also right about x87 perf being slow compared to everything else, but it’s still faster than anything you can do in software (addition especially does not work interact nicely) due to the GRS tracking a software impl needs to do through many bitewise operations.
> Even lowering the precision of the x87 unit isn’t sufficient as that only reduces the precision of the mantissa not the exponent.
I don't know what you mean by "isn't sufficient". To be clear, I'm speaking from experience emulating x86 games on low resource arm devices, where I had success emulating x87 in lower precision.
For QEMU, IMO the bigger performance issue is that it doesn't natively JIT _any_ FPU or vector instructions, and the indirect memory mapping hurts general performance quite a bit too.
The x87 unit has control bits that you control (shocking!) behaviour of the unit, one of the things you can control is the precision it will operate at operate. People think that if the x87 unit is in the lower precision modes it's possible to simply use the common 32 or 64 bit FPUs, but the x87's 32/64 bit modes only impact the mantissa, not the the exponent so they reduced precision modes are still not interchangeable with fp32 or fp64.
FEX can be quite fast at some workloads, but was slower than QEMU for others, and had some glitches for me. I ended up porting my app to arm64 Linux for Linux dev on M1 rather than continue to slog through the issues I had with emulation on Linux.
> I couldn’t include FEX in the bench as it’s not compatible with the 16k page actualy used on Asahi/M1.
FEX ran fine for me in a Parallels VM on M1.
The article discusses differences in floating point handling and GPU passthrough, but I don't think the 7z benchmark uses either of those.
The problem as far as I can infer is that for a binary translator the translator is given a bunch of context a full VM can’t have (what random clump of bytes is an executable, etc)