"With added improvements in QEMU from ICT, Loongson-3 achieves an average of 70% the performance of executing native binaries when running x86 binaries from nine benchmarks."
Loongson is MIPS which executes x86 binaries at 70% speed of native execution. Using QEMU.
QEMU is not unperformant.
Furthermore, having worked with QEMU in several binary translation and emulation projects, I cannot say that it is designed with performance as its first priority–TCG's number one consideration seems to be portability and ease of supporting new architectures. If you look at TCG-generated code it's pretty "stupid"; almost no optimizations are applied at all. This isn't necessarily something I fault QEMU for, it's just clearly not a priority. Or hasn't been, I guess; but it seems like this might be changing and I'm very interested to seeing where it goes.
So a compiler from one architecture to another which is as easy to write as an interpreter (but without interpreter overhead), is independent of the target architecture - and it just works. I remember being blown away by the elegance.
https://en.wikipedia.org/wiki/Partial_evaluation
https://labs.oracle.com/pls/apex/f?p=LABS:0::APPLICATION_PRO...
MIPS needs about 20 to 40 cycles for page fault handling. For 500MHz MIPS (Verilog/VHDL implementation synthesized to 90nm - yes, I am that old) it translates to latency from 1us to 2us. [1] shows that typical x86 page fault latency is ~7.8us, about at least twice as much. [2] (best I've found, sorry) speculates that functions compute x86 status codes and allow 80-bit long double operations natively.
[1] https://makedist.com/posts/2016/10/10/measuring-userfaultfd-...
[2] https://news.ycombinator.com/item?id=15543718
Excuse my rant, but MIPS has just one mistake - delay branch slot. It complicates implementation immensely. Other than that, it is a fascinating architecture. They claimed that their CPU can be programmed using regular instructions as effective as others do with microcode and thus MIPS does not need that and my experience had shown that.
Which isn't to say that you can't do better than 3x-slower under any circumstances, just that if you wanted performance you'd probably be better off starting from scratch with an emulator design that cared about performance and which was really clear about its use-cases -- eg "this is user-space only, not system emulation" and "this is only this very small set of host and guest architectures". QEMU does a lot of different things in one codebase, which makes it cumbersome to change anything and hard to put in optimisations or simplifications which might be valid for a specific host/target combination but not more widely.
See box86:
One workaround for that restriction is to use Cosmopolitan Libc: https://justine.lol/cosmopolitan/index.html It bundles compiler runtime libraries that are fast and permissively licensed. It lets you keep using GCC without facing restrictions on things like code morphing.
https://www.hpl.hp.com/techreports/1999/HPL-1999-78.html
Note that CPUs and compiler optimization have improved over the last 20 years and these results may not still hold.
It’d be interesting to see the performance of QEMU if it used WASM as an intermediate representation, and then fed that into WASM engine so it could take advantage of all the browser’s optimizations.