Emulating an emulator inside itself. Meet Blink
hiro.codes
hiro.codes
Another great example of this tinier is better phenomenom, would be v8 vs. quickjs. Fabrice Bellard singlehandedly wrote a JavaScript interpreter that runs the Test262 suite something like 20x faster Google's flagship V8 software, because once again, tests are ephemeral. It's amazing how much quicker QuickJS is. But if you wanted to do something like write a JS MPEG decoder to show television advertisements without a <video> tag then v8 is going to be faster, since it's a locomotive.
Fabrice Bellard wrote Qemu too. But I suspect his Tiny Code Generator has gotten a lot heftier over the years as so many people everywhere contributed to it. I really want to examine his original source code, since I'd imagine what he originally did probably looked a lot more like Blink than it looks like modern Qemu.
It seems to me like browsers would benefit from running most code in QuickJS, and then spinning up V8 only for those rare cases of long-running JS?
No he didn't.
jonah_hill_fuck_me_right?.jpg
(We don't agree on the appropriateness of the description that he "was involved", however. That sounds like an attempt to understate/diminish.)
Something is wrong here. How did you test this? QuickJS might start up faster on very small testcases but V8 is not that slow; it needs to have very low latency on a webpage too. Did you run a debug build or something?
QEMU's code generator is actually pretty fast and shouldn't really be expensive. It's a handful of passes that are run on individual basic blocks, certainly not optimal when a lot of code runs once as is the case for a very short compile but it's nothing like v8.
I suspect an even more silly reason—startup time might even be the biggest factor, because I think qemu-user's startup has never been optimized. I assume both QEMU and blink binaries are statically linked (or both dynamically linked, alternatively)?
Anyhow these theories should be pretty easy to disprove just by compiling something larger than hello world, so I will do it in case there's some low-hanging fruit left.
Even if you remain in the Unix world (so macOS but not Windows), system-call–level translation will let some of the differences between operating systems transpire to the program; and it will cause failures if a program has #ifdefs to distinguish between Linux and Darwin but then sees Darwin behavior under Linux.
On Windows, you would basically be reimplementing cygwin—imagine what a mess it would be for a Linux program to see Windows paths with drives and backslashes.
Even though blink can run any Linux program in principle, the main idea is to use it for the author's "cosmopolitan libc", to run programs that are essentially scripts compiled to machine code. So you'd write these programs portably but still be able to ship a single binary.
QEMU has all of those but only within a basic block, so it does not do any complicated data flow analysis. It's what would be considered a baseline JIT in the JavaScript world. Rosetta is similar too as far as I understand.
I suspect that the reason why blink is faster in this experiment is not related to code generation, as I mentioned in another comment. Looking again at the screenshot, the 28 context switches (vs 0 for blink) might be a clue as well.
While I have your attention, may I ask if you've considered saving your JIT's output to a file, that can be rapidly loaded if GCC is launched a second time? Ephemeral commands tend to be executed independently many times, which would seem to favor an AOT approach. But I don't see any reason why a JIT can't persist its output to gain the advantages of AOT. I'd call it EOT.
I don't work that much on QEMU TCG actually, but it would be possible to do so. In fact recent (unrelated!) changes to QEMU might even make it possible to preserve ASLR with such a first-execution JIT compilation.
That said I have rarely seen code generation in QEMU's profiles. In the end what matters is real world performance and I need to redo the test myself to understand what's going on and whether your benchmark is representative of e.g. building a small but nontrivial program (let's say blink itself) with both QEMU and blink. In that case there would be repeated cold-start recompilation, but also the compiled code would run at least once per function so QEMU would have an edge.
To be honest blink is probably more like a bicycle than a sports car. It will start faster than a locomotive, but the cruise speed is definitely lower.
Top speed of most trains is below that of a Tesla let alone super cars. The trains that can go significantly faster (but not really hundreds of miles per hour faster - Maglevs - are capable of similar acceleration). This analogy is not the greatest for the point being conveyed.
I suspect if you are coming from the perspective of a US rail user then yes, they aren't exactly known for high speed train travel.
For simple workloads it can even be faster than native unless you dynamically load something that uses bigger pages for your native program, eg. https://easyperf.net/blog/2022/09/01/Utilizing-Huge-Pages-Fo...
And none of that accounts for the increased context switch time.
What context switch time? It takes 5 micros to enter and leave the guest. The rest is just "workload".
The point is: KVM is native speed if you never have to leave. I don't need to prove this for anyone to understand it has to be true.
The guest has it's own page tables above the nested guest phys->host phys tables.
> What context switch time? It takes 5 micros to enter and leave the guest. The rest is just "workload".
And then the kernel doesn't know what to do with nearly every guest exit on KVM, so then you trap out to host user space, which then probably can't do much without the host kernel so you transition back to kernel space to actually perform whatever IO is needed, then back to host user, then back to host kernel to restart the guest, then back from host kernel to guest. So six total context swaps on a good day guest->host_kern->host_user->host_kern->host_user->host_kern->guest.
The topic of this thread is about Blink, which happens to be a userspace emulator. Hence my comment.
> Blink is now outperforming Qemu by 13% when emulating GCC.
Nice work. But isn't QEMU notoriously slow?
$ o//blink/blink /bin/ls
error: unsupported executable; we need:
- flat executables (.bin files)
- actually portable executables (MZqFpD/jartsr)
- statically-linked x86_64-linux elf executables
Blink really isn't intended to run the binaries that come with your distro, because you're already able to run them. Blink is intended to let you transplant x86-Linux executables onto non x86-Linux systems.For example, I like to build programs on my x86 Alpine Linux machine and then scp them onto my M1 MacBook, Raspberry Pi, FreeBSD, and Windows machines. Blink lets me run them once I do that. Copying program files only makes sense if the program files are static. Linux distro executables can't be distributed to somewhere other than the distros that created them, unless a tool like Docker is used which recreates the whole distro.
Blink is for people who want to be able to distribute Linux software without having to be in the Linux Distributor business too.
At least I feel I'm regularly installing things from people that are not in the Linux Distributor business and are not huge go binaries either (though admittedly, those seems to be trendy as well :) )
It's a nice test also and can be useful for debugging.
I've only ever really skimmed the TCG source code but it wouldn't surprise me if a new-er JIT could smack it's arse given that with these old C codebases (it's probably one of Bellard's few flaws) it's pretty hard to actually make true architectural changes.
The Java/script (I think more Javascript but I'm hedging my bets by including jvms too) JITs are probably the cutting edge but I'd imagine still quite beatable for a few cases.
Looks like it itself is not yet able to be compiled with Cosmopolitan Libc (though it emulates programs compiled with it) but it's planned - very cool!
I'm not a windows user, but super interested in using Blink to ship pre-compiled binaries as part of various Bazel rule sets.
My favorite FAQ
With memory safety turned on, could Blink be used as an alternative to WebAssembly?
I think the user experience of cli debuggers is generally somewhat dreadful when compared to their gui cousins -- they seem to display a much narrower view of what's going on. Could the big blinkenlights debugger view be useful outside of blink itself?
* https://github.com/pwndbg/pwndbg * https://github.com/longld/peda * https://github.com/hugsy/gef
Anyone know of any other use cases they have in mind?
I always hear tech spec rumors but never about anything I would want to do with this type of thing… outside say gaming?
I recommend obtaining Python from the Comsopolitan mono repo:
git clone https://github.com/jart/cosmopolitan
make -j8 o//third_party/python/python.com
That'll easily give you a static Python 3.6 binary that can be run under Blink.[1]: https://blink.sh/