The Evolution of the QEMU Translator
linaro.org
linaro.org
The main idea as I understood/remember it was roughly this:
UEFI offers an abstract hardware interface to the OS, but to actually talk to the hardware it also needs drivers. To solve this issue, UEFI defines a driver interface and loads driver blobs directly from a ROM on the device itself. However, if you plug an off-the-shelf PCI card into an ARM server, this doesn't work, because the driver blob is x86 code. For Linux this doesn't matter, since it has its own drivers, but e.g. Grub doesn't and if the PCI device in question is a graphics card, you are flying blind until the kernel is up.
So they patched their version of EDK2 to map the option ROM blobs as not-executable and when UEFI/the bootloader tries to call into the option ROM blob, the page fault handler they installed traps that, extracts the call arguments and runs the blob through TCG. They use the same trick to catch calls from the ROM blob back into their ARM UEFI binary, extract the arguments again, do the call, convert the return values and jump back.
It sounds like a crazy hack, but apparently worked well enough that they could plug an Nvidia card into an ARM server and it would display the Grub boot splash screen and early printk messages.
[1] https://osseu17.sched.com/event/ByIv/qemu-in-uefi-alexander-...
http://landley.net/hg/qcc/file/tip/todo/todo.txt
It seems like a pretty cool idea, so I'm hopeful that eventually he'll have the chance to hack on it and get it working.
Maybe set up a chat and see what happens?
"With added improvements in QEMU from ICT, Loongson-3 achieves an average of 70% the performance of executing native binaries when running x86 binaries from nine benchmarks."
Loongson is MIPS which executes x86 binaries at 70% speed of native execution. Using QEMU.
QEMU is not unperformant.
Furthermore, having worked with QEMU in several binary translation and emulation projects, I cannot say that it is designed with performance as its first priority–TCG's number one consideration seems to be portability and ease of supporting new architectures. If you look at TCG-generated code it's pretty "stupid"; almost no optimizations are applied at all. This isn't necessarily something I fault QEMU for, it's just clearly not a priority. Or hasn't been, I guess; but it seems like this might be changing and I'm very interested to seeing where it goes.
So a compiler from one architecture to another which is as easy to write as an interpreter (but without interpreter overhead), is independent of the target architecture - and it just works. I remember being blown away by the elegance.
https://en.wikipedia.org/wiki/Partial_evaluation
https://labs.oracle.com/pls/apex/f?p=LABS:0::APPLICATION_PRO...
MIPS needs about 20 to 40 cycles for page fault handling. For 500MHz MIPS (Verilog/VHDL implementation synthesized to 90nm - yes, I am that old) it translates to latency from 1us to 2us. [1] shows that typical x86 page fault latency is ~7.8us, about at least twice as much. [2] (best I've found, sorry) speculates that functions compute x86 status codes and allow 80-bit long double operations natively.
[1] https://makedist.com/posts/2016/10/10/measuring-userfaultfd-...
[2] https://news.ycombinator.com/item?id=15543718
Excuse my rant, but MIPS has just one mistake - delay branch slot. It complicates implementation immensely. Other than that, it is a fascinating architecture. They claimed that their CPU can be programmed using regular instructions as effective as others do with microcode and thus MIPS does not need that and my experience had shown that.
One workaround for that restriction is to use Cosmopolitan Libc: https://justine.lol/cosmopolitan/index.html It bundles compiler runtime libraries that are fast and permissively licensed. It lets you keep using GCC without facing restrictions on things like code morphing.
Which isn't to say that you can't do better than 3x-slower under any circumstances, just that if you wanted performance you'd probably be better off starting from scratch with an emulator design that cared about performance and which was really clear about its use-cases -- eg "this is user-space only, not system emulation" and "this is only this very small set of host and guest architectures". QEMU does a lot of different things in one codebase, which makes it cumbersome to change anything and hard to put in optimisations or simplifications which might be valid for a specific host/target combination but not more widely.
See box86:
It’d be interesting to see the performance of QEMU if it used WASM as an intermediate representation, and then fed that into WASM engine so it could take advantage of all the browser’s optimizations.
https://www.hpl.hp.com/techreports/1999/HPL-1999-78.html
Note that CPUs and compiler optimization have improved over the last 20 years and these results may not still hold.
Perhaps not Linux-complex, but certainly close. So it's somewhat impressive that it has come so far when the Developer docs states: QEMU does not have a high level design description document - only the source code tells the full story.
But it can run Windows 10 on Mac :) https://forums.macrumors.com/threads/success-virtualize-wind...