V8: A Tale of TurboFan
benediktmeurer.de
benediktmeurer.de
In the years since I've seen Android go from interpreter, to interpreter+JIT, to AOT, to JIT+AOT, to interpreted + JIT then AOT at night. V8 has gone from only JIT (compile on first use) to multiple JITs to now, interpreter+single JIT. MS CLR still doesn't have any interpreter and is fully AOT or JIT depending on mode. HotSpot started interpreted, then gained a parallel fast JIT, then gained a parallel optimising JIT, then went to a tiered mechanism where code can be interpreted, compiled and recompiled multiple times before the system stabilises at peak performance.
Looking back, it's apparent that Android's and indeed Java's initial design was really quite insightful. A tight bytecode interpreter isn't quite as awful as it sounds, especially given how far CPU core execution speed has raced ahead of memory and cache availability. If you can fit an interpreter almost entirely in icache, and pack a ton of logic into dense bytecode, and if you can use spare cores that would otherwise be idle (in desktop/mobile scenarios) to do profile guided optimisation, you can end up utilising machine resources more effectively than it might otherwise appear.
That is why when you research old papers about compilers, Assembly opcodes and Bytecode are often intermixed.
It was common for the OS, language runtime to work with bytecodes, with the execution being done at microcode level. What nowadays would be better done as FPGA.
On IBM i (aka OS/400), there is a kernel JIT and all languages target bytecode, including C. Also JVM bytecodes get translated into IBM I ones. It was implemented in PL/M, nowadays with many modules ported to C++. And if you want to generate actual machine code, you need the Metal C compilers.
There are plenty of other examples.
Like with containers, it is just recycling mainframe success stories.
With Android there are hundreds or thousands of combinations of various bits of hardware and drivers, with various levels of compatibility. Device manufacturers have a terrific advantage in being able to produce hardware that is more specialised for the purpose, but there are trade-offs. Compiling all possible variations AOT away from the end device is going to be difficult.
I guess nothing would stop the Play Store from AOT compiling binaries just for users with fast connections though.
WP 8 and 8.1 use MDIL, Machine Dependent Intermediate Language, basically machine code with jump targets kept as symbolic links, replaced at installation time.
WP 10 uses .NET Native.
I can gladly point out the respective Microsoft documentation, BUILD and Channel 9 presentations.
And then you'd still have to fall back to on-device compilation if Play Store doesn't know what the device is anyway (aka, all custom ROMs), or when you take an OTA (otherwise you'd need to re-download all installed apps for the new OTA image).
Plus, the dex bytecode is considerably smaller than the AOT result, so wire transmission costs are cheaper doing on-device compilation - call it a form of compression.
So it's a complex topic.
Windows Phone .NET runtime also uses tracing GC, although the underlying UWP APIs are COM based.
The specs of the various devices have changed over time. Android scales better: you can give it very little RAM, or a lot, and it'll make best use of it by e.g. keeping more background apps loaded at once. iOS was at least historically much less flexible about this: boosting RAM was not worth as much as in the Android space because app devs would still target older devices for quite a while and so the additional RAM would go unused, and the OS wasn't capable of using the spare for much due to the lack of multi-tasking.
Eventually Apple implemented Android-style task switching, so I don't know if that's still true. I haven't done any mobile dev for years. But I also think at some point Apple realised nobody who buys an iPhone actually cares about whether they're getting value for money or what the specs are, so they just stopped competing on that area. I mean they have never reduced the price of the iPhone once, right? Despite the huge fall in underlying component prices over the years. They could ship a device with 128mb of RAM and if it did the same thing as the iPhone 4 people would still buy it, simply because they see themselves as iPhone users and not "smartphone users.
ObjC, as typically written (i.e. more object-oriented than C), is also pointer and allocation heavy. so that might make compaction more of a benefit for Android. I wonder what proportion of memory in a typical Android application process is on the C heap rather than the garbage-collected, compacted heap. For example, where does image data usually end up?
To throw in another complication, are there any significant problems that come with layering a garbage-collected runtime on top of a high-level framework based on C heap allocation and reference counting? That's what Xamarin, React Native, and AOT Java runtimes (e.g. Multi OS Engine) do on iOS, and what .NET (even .NET Native) does on UWP. Or how about two garbage-collected runtimes in one process, e.g. Xamarin and React Native on Android?
Does Android use more memory than iOS?
The difference from Apple and Microsoft's bytecode solutions versus Google, is that the AOT compilation of bytecode takes place at the store, instead of the device.
dalvik's GC was a big problem for smooth UIs, though.
By the way, I know this is a tenuous association, but to me, the name TurboFan makes me think of CPU-hungry code that cranks up the fan(s) on a user's laptop.
While performance is not one of the main concerns for the library (they are simplicity, "batteries-on" and user experience in general) it is nice to see that my hunch was right; V8 would provide a 500% performance boost to Promises.
This could be extremely useful. I've been using express in production for years, and though setting up a new server is not fundamentally difficult, it does involve a lot of common boilerplate (cookie-parser et al) that I think a lot of people already solve with a boilerplate template. This could be a cleaner, more fully-featured approach.
Is this more or less it, or is there more to it / future aims to do more? I also wonder - did you talk to the express developers about making direct contributions to provide solutions to these problems from within the express library itself, rather than externally via a wrapper library?
I'll keep an eye on the development of this project :-)
The main differences (for the initial or a later release) are:
* A lot of functionality out-of-the-box. Just install it and get to work.
* Websockets as first-class citizens: since they are one of the major advantages of Node.js in itself it makes sense they are trivial to use.
* Error handling: intercept some messages from Node.js and provide a more human-readable version.
* Promise-based.
There's also a single parameter for middleware instead of a [err], req, res, next parameters, since Promises work really well with a single parameter. You might think that this comes from http://koajs.com/ , but only the name comes from there, I used to call it inst for instance until I found a better name, which I found in Koa's ctx.
Oh, I talk about all of this in here: https://serverjs.io/about
100x smaller, multiple times faster and did it right from the beginning.
That said LuaJIT is an impressive piece of software. Would love to know what Mike Pall is up to now.
The Lua ecosystem is very different than almost any other; It's primary use case, probably 95% of projects using it or more, is being embedded into another project for scripting/configuration/control. As such, it is common for projects to pick a version of Lua and stick with it, rather than upgrade to the latest-and-greatest-with-slight-to-major-incompatibilities. Pall liked 5.2, and thought 5.3 didn't offer enough to break compatibility.
You sure about that? In the alioth benchmarks, v8 wipes the floor with lua: http://benchmarksgame.alioth.debian.org/u64q/compare.php?lan...
"If you're interested in something not shown on the benchmarks game website then please take the program source code and the measurement scripts and publish your own measurements."
Like this guy did -- https://pybenchmarks.org/u64q/which-programs-are-fastest.php
This is an unsubstantiated claim. If LuaJIT was indeed multiple times faster on arbitrary code, everybody would be just running JS on something like LuaJIT in browsers :)
LuaJIT does spectacular job on loops, especially if you limit yourself to number crunching and FFI.
But if you code is polymorphic, makes heavy use of normal OOP and does not decompose into a graph of biased loops and linear traces V8 will pull ahead due to its method-compiler nature which is not susceptible to tracing pitfalls.
Here is an interesting example: there is an issue[1] "Metatable/__index specialization" in LuaJIT repository filed by Mike Pall himself, it's about adding infrastructure which would allow traces to make optimistic assumptions about metatable constness, because it would greatly improve the quality of traces produced from OOP heavy code. (Also this is precisely the limitation that is already imposed on FFI metatypes to allow JIT generate better code). The issue is still unimplemented in LuaJIT... However this is something that V8 could do for several years now, because it is extremely important for the kind of code people are writing in the real world.
What were the weaknesses of the Sea of Nodes? Backwards data flow analysis and control flow sensitive analysis being hard?
Changes to the CFG are hard because of the Phi nodes, which have to be adapted in lockstep with CFG changes. However, this is probably inherent to SSA form, but not sea of nodes.
You want a good graph visualization instead of text-based output, because you should actually to see the "non-order of the sea". Text-output with an implicit order can hide issues.
Personally, I believe sea of nodes is better then others, like SSA form is better than non-SSA. There is nothing which makes it inherently more powerful, but it feels more elegant. Unfortunately, there is no objective comparison and anecdotes are apples vs oranges. Similarly, how would you compare object-oriented with functional programming?
I loved your paper about PBQP register allocation, by the way!
While I've often felt many of the pain points in the article and they were never explained anywhere (that I saw anyway).
I'd often profile/benchmark something, make sure it's fast enough to use in our performance critical section of the code, only to find that once in the application I'd only get a fraction of the speed I was expecting.
I would then hop over to using --trace-opt only to find that functions were getting deoptimized or never optimized in the first place and I'd start playing the game of trying things here and there to get it to cooperate. And in some cases --trace-opt wouldn't tell me anything that I could usefully understand yet my code would still be slow.
Here's to hoping that turbofan clears up a lot of these weird cases!
And slightly off topic, but what are the plans for the dart VM? Is it going to end up using TurboFan or will it stay with Crankshaft?
Dart VM is a code base independent of V8, so V8's plans to adopt one or another compiler have no implications on Dart VM.
Dart VM's IR is closer to Crankshaft, but it supports full language unlike Crankshaft (so in this sense it is closer to TurboFan).
We currently have no plans to radically rework the compilation pipeline, because we see no need for that.
I've never managed to find a good source on the V8 internals and how to target these optimizations (the "rules for the arguments object" the author alludes to).
Any recommendations?
1. https://github.com/petkaantonov/bluebird/wiki/Optimization-k...
And current 5.7 issues seem to live here: https://github.com/nodejs/v8/issues
https://github.com/nodejs/node/commit/7a77daf24344db7942e34c...