V8 Sparkplug – A non-optimizing JavaScript compiler
v8.dev
v8.dev
> Enter Sparkplug: our new non-optimising JavaScript compiler we’re releasing with V8 v9.1, which nestles between the Ignition interpreter and the TurboFan optimising compiler.
This surprised me. Didn't V8 have a 'fast JIT' alongside an optimizing JIT as early as 2013? It was called full-codegen. Did they remove it?
https://wingolog.org/archives/2013/04/18/inside-full-codegen...
Did you see the image at the bottom of the design document? That explains it.
[1]:https://docs.google.com/document/d/13c-xXmFOMcpUQNqo66XWQt3u...
For example: * What is bottom? Did it means one of the pictures in the bottom page. Or the absolute bottom.
* What if the picture was not at the bottom, instead, it was followed by large chunk of texts?
* What deisng doc? Was it mentioned in the parent post, or was it a "well known" design doc?
* Design doc for which topic? Turbofan/Craneshaft/etc? (This is nitpicking, but still happen often because one might found this as the first sight, not following the thread. Again, reading answers for engineers are also a skill, quite different from reading, for example, science fiction, another day another topic)
* What does the figure explain? What if the figure shows something that not match my concept of the question?
In the end, answering engineering question, should be, and have to be, stating accurate facts. If and how that actually answers the question, is more depending on the readers and questioner's continued communication.
Edit: To make sure I was not leaving wrong impression, those questions I listed above was encountered everyday for the answers I received. It's not nitpicking. You can read this thread for an example how answers without concrete information being inappropriate engineering answers.
BTW, I view this as not about communication style, but communication efficiency.
They removed it a while back after they introduced the ignition interpreter: v8 used to only have compilers (full-codegen and crankshaft, the latter being the jit optimising compiler, full-codegen was not a jit).
One particularity of the pipeline was that since there was no interpreter both compilers worked from the source, which crankshaft had to keep around and re-parse.
In 2013 they started working on a new optimising jit (turbofan), but for a while it didn’t really improve things (though the failure modes were often better than crankshaft’s).
Then some engineers realised they could use part of turbofan for an interpreter, which would allow load sharing (amongst other things turbofan would jit ignition’s bytecode instead of having to re-parse from source), that brought in the improvements they were looking for, but then full-codegen was kinda sitting in the middle awkwardly with no integration, its low-end eaten by ignition and with no real high end. So it was dropped alongside crankshaft in v8 5.9.
Now they’re reintroducing a non-optimising compiler, but one that’s integrated with the ignition-turbofan pipeline, and actually a jit.
By what definition?
With an interpreter, the JavaScript source is compiled to some bytecode, then the bytecode is executed in a loop, instruction by instruction.
With a compiler (not jit), the whole bytecode would be further fully compiled to some C apis before being executed, making it effectively an "ahead of time" compilation.
The catch is : since the code never was executed, there is very little static information to work with, so the compiled output is often just a giant unrolled interpret loop. The compiled version can be like 10% faster on very hot JavaScript loops compared to the interpreted version.
This hinges at least on the question of what you mean by "the whole bytecode". If the source is made up of N functions, and you compile all N functions ("the whole bytecode" of the entire program) to machine code before you know if any of them will be executed: Yes, I'd agree that that's a form of ahead of time compilation. Is this what full-codegen used to do?
On the other hand, if the source is made up of N functions, and you only compile "the whole bytecode" of each of them to machine code one by one, only those that will actually be executed, and only when you are preparing to actually execute them: That's JIT compilation. Even if it is the very first execution (i.e., you have never even interpreted), even if the compilation doesn't use fancy profiling or sophisticated code generation or whatnot.
> since the code never was executed, there is very little static information to work with, so the compiled output is often just a giant unrolled interpret loop
I think you mean "dynamic" information. And "the compiled output is often just a giant unrolled interpret loop" is pretty much the definition of what a baseline JIT compiler produces.
Yes. Afaik full-codegen had no way to discover that.
> and the "X" is never "FullCodeGen was not a JIT". Which suggests to me that the people closest to development of these things don't think of them in the same terms as you do, which is why I'm trying to understand your terms.
If you look at the langage they used, they’d always call full-codegen a compiler or baseline compiler. Not a jit. Because there was nothing to gather time-of-use information from. And full-codegen didn’t keep parse information around which is why crankshaft had to separately re-parse everything, full-codegen could not hand it the source or bytecode of functions to optimise.
The big difference is that Sparkplug compiles from bytecode, not from source, and thus the bytecode stays the source of truth for the program. Back in the FCG days, the optimising compiler had to re-parse the source code to AST, and compile from there - even worse, to be able to deoptimise back to FCG, it had to kind of "replay" FCG compilation to get the deopted stack frame right.
We did actually have a configuration (never released) that had Ignition, FCG, _and_ TurboFan, and boy was it a mess...
Do you mean it had technical debt that was never resolved, or do you mean 3 tiers is always too many?
I believe WebKit's JavaScript engine currently has 4 tiers, [0] although that includes an interpreter as the first tier.
[0] https://arstechnica.com/information-technology/2014/05/apple...
The exact tiers have changed, but it’s still 4 of them: an interpreter (LLInt), the Baseline JIT, and two optimizing JITs
Three tiers is perfectly fine if the complexity is handled well, I could even see us adopting a fourth tier at some point in the future (not dissimilar indeed to JSC), most likely between Sparkplug and TurboFan.
It'd be nice to have some simple benchmarks that directly compare ignition, sparkplug and turbofan (and full-codegen too). And also time spent compiling.
It looks somewhat similar to context threading as I remember it:
http://www.cs.toronto.edu/~matz/pubs/demkea_context.pdf
but it's been a while...
Ha-ha, the stacks grow downwards button is priceless :-D
> Isn’t this just FullCodeGen?
> Kind of! It’s FullCodeGen for modern V8, and that’s not such a bad thing. This brings the performance benefits of FCG back to V8, but
I like the reasoning in the sibling comment:
> So if you want to understand what other people are saying about "compilers" you should keep in mind that the standard definition includes what you are calling "transpilers".
Really, the main disadvantage of saying "transpile" is that it implies you don't know what "compile" means. In a context where that wasn't an issue -- e.g., if I heard it from one of my coworkers, who certainly know what a compiler is -- I would just think they liked the neologism.
Isn't bytecode an intermediate representation?
https://en.wikipedia.org/wiki/Intermediate_representation
Am I just being pedantic?
In this case, if they didn't have an interpreter, they'd probably use an abstract syntax tree or control flow graph as their common source of truth for their two compilers. So, the primary reason the bytecode exists is to be interpreted by the interpreter and therefore isn't intermediate.
(LLVM IR is an unusual exception therefore. Please don't 'well actually' me I'm trying to give a useful practical explanation.)
Bytecode is usually a buffer of high level instructions that is compact and fast to read (for the machine). Good for interpreting and holding the semantics of a program in a format that is more efficient than the source code, but it is usually not a good format for running compiler passes.
To me, an intermediate representation is what the compiler translates the input to other than machine code. Sparkplug goes straight from its input (bytecode) to machine code.
So you’re not being too pedantic. It’s just that V8 has two compilers now, one that has an IR of its own (turbofan), and one that doesn’t (sparkplug). Lack of IR is one of the things (maybe the most important thing) that makes Sparkplug special. If you define IR in a way that means that bytecode is an IR then I don’t know how to easily articulate what it is that makes Sparkplug special.
I could see them splitting Ignition into a bytecode compiler and separate interpreter and give them separate names.
From what's in the post, we can't tell if the interpreter is 128× slower than optimized code while Sparkplug code is only 32× slower, or the interpreter is 32× slower and Sparkplug code is 16× slower, or the interpreter is 32× slower and Sparkplug code is 4× slower (which is my guess as to the reality of the situation.) I'm more interested in this is-it-4×-or-32× question than the possibility that microbenchmarks might misrepresent a 32× slowdown as being a 48× slowdown or a 24× slowdown.
https://docs.google.com/document/d/13c-xXmFOMcpUQNqo66XWQt3u... does offer more performance numbers, but they're for "an early prototype", and more importantly they don't break out the compilation speed vs. the execution speed.
The description in the post makes it sound like Sparkplug output is pretty similar to direct-threaded Forth code, just a sequence of subroutine calls to other methods and implementations of primitive operations, interspersed with the occasional open-coded primitive (=== maybe) and control flow. And typically that means an overhead of around 4×, although maybe it's more with JS's aggressive implicit coercions and whatnot. It sure would have been nice to see some disassembly listings, too.
From what I understand, there are basically 3 things that provide the bulk of performance gains in compilers: inline caches, register allocation, and inlining. The latter two don't really apply to interpreters or baseline compilers.
Compilers do many optimizations beyond those, but generally only for incremental improvements.
I think constant folding, specialization, and loop hoisting (especially of bounds checks) commonly produce more than merely incremental improvements. If you don't have some kind of inlining you don't get those (usually) but inlining by itself is usually only a minor improvement.
ICs steal a lot of specialization's thunder. (Alternatively, since the point of ICs is specialization, you could say it's the other way around!)
Edited to add: Which is a very good result! Congrats to the V8 team for landing this!
Compile time is on roughly the same order of magnitude as Ignition compilation (just AST to bytecode, so excluding parsing), and roughly two to three orders of magnitude faster that TurboFan. The relative performance to the interpreter, as the sibling comments point out, varies _wildly_ by workload, but around 4x is probably a decent approximation for something not entirely dominated by property loads.
#define __ basm_.
Btw. the world of browsers currently is totally crazy. You have people dropping to HTML canvas and own rendering because manipulating the DOM is so slow for anything serious. There are bugs in CSS and SVG all over the place, you cannot really read more in depth documentation how to program the web for performance in this space. (If you have great, very in depth documentation, please link it!) The situation is so bad, Google teams rather switch to their own rendering in HTML canvas instead of talking to their peers and letting them fix the DOM manipulation for everybody. Actually, Chrome/ Chromium feel progressively worse with each release. There are always new regressions and there don't seem to be any real improvements in efficiency. The result is, we are busy working around idiotic APIs and bugs like 90% of the time. Yeah, OrgPad is not a trivial app, but we have something like 1000x better hardware than in the 90s and some of the interactions feel much slower than the games we had back then.
Does anybody else have a similar experience or some actually helpful resources we possibly overlooked?
It's in progress but we cover the core architecture of a modern browser in some depth. It's funny you mention CSS recomputation... I have a student working on CSS and layout recomputations (with the Google Chrome team, who are fantastic) right now. It's a challenging and exciting area.
Most interactions using a 2018 iPad Pro w/ Safari take seconds to execute. This device can handle anything I throw at it except reddit.com.
The new web app is a plain regression.
That’s why I switched to teddit.net. It’s much faster, resembles the old Reddit layout and supports dark mode.
I have pretty much the same thoughts about discourse, which I loathe.
The sad thing is, the front end stack they use (React/Redux) is great. But it's an abysmal example of the tech. I turned on re-renders in the React devtools, and this is what it looks like: https://i.imgur.com/Sm1U24M.gifv
Makes me sad.
After clicking the link to the article, I hit the back button back to HN and a Chrome popup display "Click Install V8 to get back to this site easily" or something of that nature. It was definitely a chrome popup triggered by leaving the site, since the popup was on top of the bookmarks toolbar.
Anyone notice that happening as of the latest chrome release? I've seen this a few times, and it always catches me off guard.
Sparkplug is designed specifically as a baseline compiler that does not have an IR, in order for maximum compile speed. While the idea of two compilers with different compile time/speed tradeoffs is similar, the constants involved are pretty different.
TurboFan is similar to C2, but Sparkplug is not similar to C1. They don't compete, but cooperate.
It‘s really a pity how Qt applications such as kmail became less portable when the underlying web framework was switched from WebKit to Chromium/Blink.
Or take the existing Webpack CLI which is written in Node, what kind of improvements can we expect there, if any?
The distinction between compiled and interpreted languages is also a bit fuzzy to begin with.
I wonder how will it compare to JSC.
My guess is that all of the main JS engines will eventually evolve into crabs.
Initially v8 only had a baseline compiler and an optimising jit (in fact that was very much the originality of v8 at release).
A few years back they switched to an interpreter + optimising jit (ignition / turbofan), then they added the mid-tier turboprop which is turbofan with the most expensive optimisations dropped.
* Rewrite those bits of javascript in C++, and patch them in whenever they are seen byte-for-byte in the wild.
* Have both a fuzzer and a webcrawler verify the two implementations never differ.
Obviously it's vital the C++ and javascript implementations behave identically - so probably best to have a bunch of checks that the input is in the exact form expected, and fall back to the javascript version if the C++ version can't be guaranteed to behave the same.
The C++ versions of say the globally most popular 25 MB of javascript source code could be shipped with Chrome, or as a seperate download-on-first-use component.
Saving a bit of JavaScript execution time on the most costly JS snippets would have translated into real money savings in terms of both capital and operating expenses.
Especially since there were tons of pages where the most popular JS frameworks were being loaded and executed once, I thought about dumping SpikerMonkey bytecode (or toward the end, v8's first-pass x86 machine code) to disk/BigTable, or perhaps C++ rewrites of several of the most common JS frameworks.
I basically wrote a high-performance low-fidelity browser simulator that simultaneously pretended to be FireFox and IE (so that it would mostly work on IE-only pages and also mostly work on FF-only pages). There was some appeal to the idea of tweaking the JS frameworks to fit my simulator rather than tweaking the simulator for all these little corner cases. (I remember having to improve simulation fidelity for Saab's unified landing page for several Arab markets.)
The thing is, a high-fidelity rewrite of a small number of JavaScript farmeworks into C++ is probably roughly equivalent to writing the first version of v8, and you need to throw it all away as the frameworks upgrade versions.
There's still some merit to caching the bytecode for the most popular JS frameworks, and I wouldn't be surprised if Google currently does this for executing JS in the indexing system.