Maybe you don't need Rust and WASM to speed up your JS (2018)
mrale.ph
mrale.ph
High-perf JS relies on staying on the JIT happy path, not upsetting GC, and all of this is informal trickery that's not easy to do for all JS engines. There's never any guarantee that JS that optimizes perfectly today won't hit a performance cliff in a new JIT implementation tomorrow.
I thought HN was supposed to be better.
I would say the opposite: there kind of are such guarantees, to JIT development’s detriment. Because JIT developers use the performance of known workloads under current state-of-the-art JITs as a regression-testing baseline. If their JIT underperforms the one it’s aiming to supersede on a known workload, their job isn’t done.
This means that there’s little “slack” (https://slatestarcodex.com/2020/05/12/studies-on-slack/) in JIT development — it’ll be unlikely that we’ll see a JIT released (from a major team) that’s an order-of-magnitude better at certain tasks, at the expense of being less-than-an-order-of-magnitude worse at others. (Maybe from an individual as code ancillary to a research paper, but never commercialized/operationalized.)
The biggest problem is that JIT is guided by heuristics. Observations from interpreters decide when code is advanced to higher optimization tiers, and that varies between implementations and changes relatively often. Your hottest code is also subject to JIT cache pressue and eviction policy, so you can't even rely on it staying hot across different machines using the exact same browser.
Now, different people in the industry will have different perspectives, so I'll preface this by saying that this is _my_ view on where the future of JIT compilation leads.
The order of magnitude gains were mostly plumbed with the techniques JIT engineers have brought into play over the last decade or so.
One aspect that remains relevant now is responsiveness to varying workloads, and nible reaction to polymorphic code, and finding the right balance between time spent compiling and the productivity of the optimized code that comes out the other end. There is significant work yet to be done in finding the right runtime type models that quickly and effectively distinguish between monomorphism, polymorphism, and megamorphism, and are able to respond appropriately.
The other major avenue for development, and the one that I have yet to see a lot of talk about, is the potential to move many of these techniques and insights out of the domain of language runtime optimization, and into _libraries_, and allowing developers direct API-level access to the optimization strategies that have been developed.
If you work on this stuff, you find very quickly that a huge amount of a JIT VM's backend infrastructure has nothing to with _compilation_ per se, and much more to do with the support structures that allow for the discovery of fastpaths for operations over "stable" data structures.
In Javascript, the data structure that is of most relevance is a linked list of property-value-maps (javascript objects linked together by proto chains). We use well-known heuristics about the nature of those structures (e.g. "many objects will share the _keyset_ component of the hashtable", and "the linked list of hashtables will organize itself into a tree structure"). Using that information, we factor out common bits of the data-structure "behind the scenes" to optimize object representation and capture shared structural information in hidden types, and then apply techniques (such as inline caches) to optimize on the back of that shared structure.
There's no reason that this sort of approach cannot be applied to _user specified structures_ of different sorts. Different "skeletal shapes" that are highly conserved in programs.
For me, the big promise that these technologies have brought to the fore is the possibility of effectively doing partial specialization of data structures at runtime. Reorganizing data structures under the hood of an implementation to deliver, transparently to the client developer, optimization opportunities that we simply don't even consider as possible today.
Basically why the guess structures when programmer could easily tell you the static bits and ask them to be frozen.
`Object.defineProperty(obj, "prop", { writable: true, value: ..., backbone: true });`
Then you'd need to generalize the hidden type modeler in the VM to interpret the backbone field and lift "prop" (and the value associated with it) into the hidden type representation. The important thing to remember here would be to ensure that these foldings of things into the hidden type are _transitive_. If the value of `obj.prop` itself has one or more "backbone" fields defined on it, then those too would need to be lifted into the hidden type.
Assignments to these fields would need to generate a new hidden type for the underlying object, much as new shapes are created and installed for JS objects are mutated to add new properties or change the prototype.
You'd then need to implement some mechanism of collapsing the object representations down to some inline format.
Lastly, you'd need to teach the code-generator to use the hidden type to optimize accesses through "obj.prop" (or longer property access sequences).
At the end of all of that, you then get a really beautiful optimization behaviour where:
"obj.prop.anotherBackboneProp.someRegularProp", in suitably monomorphic or lightly-polymorphic locations, becomes not 3 separately-checked property accesses, but a single type-check on "obj", followed by a access into the top-level object.
That would give a rudimentary ability for library and framework developers to define common shared, stable backbone structures for their own use. And that's fun to really think about.
There's a bunch of work to be done, and it has to be motivated by real use cases, but I'm sure they are there. The idea that this powerful technique is only applicable to the _specific_ set of data structures involved in property lookup in object inheritance chains seems an unimaginative perspective.
If you'd simply omitted the bracketed clause, perhaps your comment might have been useful!
SSC wouldn't have a niche if they could get concepts across without a bunch of obscurantist bullshit.
People criticize SSC because they're like "uh leave it to the REAL experts", but A) it's just opinions and B) the 'real experts' are unreadable, so I'll continue to ignore them while occasionally checking out astralcodexten.
Yep. As someone who played this game a few years ago, JIT implementations do change (e.g. function inlining based on function size was removed from V8, delete performance changed, number of allowed slots in hidden classes changed, etc).
Also worth noting that just because an optimization works in V8, there's no guarantee that it will also work in another JS engine.
The essay even covers that, when it mentions that matching arities had a significant impact on V8 (14% improvement) but not perceivable effect on SpiderMonkey.
But as compilers got better this become less and less of an issue. By today it has become very rare to need to write in assembly. So shouldn't history repeat itself? That JIT compilers will be so good that it becomes extremely unlikely that you'll hit the performance cliff.
What really happened is that thanks to the advances in computing power over the decades, we can now afford to skimp on performance most of the time - putting effort in optimizing only the most critical code paths - and still be able to get away with it.
But when we're talking about squeezing every drop of performance, one of the big challenges with fighting against the JIT engine is figuring out when things deoptimize. This is because there are patterns (e.g. polymorphic code paths) that are inherently impossible to optimize perfectly (because halting problem). As soon as the runtime runs into a "weird" data type, deoptimization is forced to kick in and that can destroy performance. As a developer you have to be extremely mindful of this. Type systems like Typescript help, but you can still be effectively polymorphic (e.g. large union types, `any`, etc) if you're not being mindful of the underlying machine code.
And JIT has one major downside that these benchmarks often intentionally omit in the name of "reducing noise": they usually take a while to warm up to native-like speed. I'm talking in the order of thousands of iterations running at a couple of order of magnitudes slower before JIT speeds fully kicks in.
[0] https://github.com/jart/cosmopolitan/blob/master/libc/nexgen...
On Android and UWP's case, PGO data is even shared across devices.
And IBM has a JVM commercial feature where a dedicated JIT server gets used and serves the whole cluster nodes with heavily optimized native code.
Naturally JITs coded on long nights tend to lack such engineering.
Also, compilers often can't optimize stuff which is obvious to humans. E.g. here is an example of GCC not being able to delete useless code: <https://godbolt.org/z/655KeM33Y> (it's fixed in later versions of GCC, but this version of GCC also isn't that old). Another example is Clang emitting useless code out of nowhere: <https://youtu.be/R5tBY9Zyw6o?t=1620>. So, it's not out of the ordinary for compilers to be really stupid about the assembly they generate.
This is partly a reason why virtual dom projects don't expose the ability to replace a virtual dom node constructor with inline literals despite it being a pure deterministic function.
That said, one major benefit of WASM that Javascript jits will have a hard time competing with is GC pressure. So long as your WASM lib focuses on stack allocations, it'll be real tough for a Javascript native algorithm doing the same thing to compete (particularly if there's a bunch of object/state management).
For a hot math loop doing floating point calcs (mandelbrot calc), however, I've seen javascript end up with identical performance compared to WASM. It's really pretty nuts.
A javascript JIT is never going to be able to compete with codegen from a low-level statically typed language running through an optimizing compiler. I mean, this very article contains the perfect example: by manually inlining the comparison function they got huge performance gains from a sorting function. That is child’s play for GCC or LLVM (it’s the whole point of std::sort in C++).
What I'm saying is that you'll find more often than not that it gets close enough to not matter. You'll see javascript at or below 2->3x the runtime of a C++ or Rust statically compiled implementation in most cases. That is pretty close as far as languages go. Java is right in the same range (maybe a little lower).
It’s the static-typing, not the low-level-ness, doing most of the heavy lifting in making code JITable/WPOable. You don’t need to manually inline a comparison function, if the JIT knows how to inline the larger class of thing that a comparison function happens to fall into, and if the code is amenable to that particular WPO transformation.
I would compare this to SQL: you don’t optimize a SQL query plan by dropping to some lower level where you directly do the query planner’s job for it. Instead, you just add information to the shape of your query such that it becomes amenable to the optimization the query planner knows how to do. That will almost always get you 100% of the wins it’s possible to get on the given query engine anyway, such that there’d be nothing to gain by writing a query plan yourself.
This is why, for example, javascript in Graal ends up running nearly as fast as Java in the same VM. The reason it isn't just as fast is the VM has to insert constraint checks to deopt when an assumption about the type shape is violated.
JIT is theoretically possible to figure out all that and transform your code into optimized form, but practically? we don't have Sufficiently Smart Compiler[1] yet. Usually, JIT is worse in restructure your data layout than figure out which part to inline.
That said, yes, the runtime information provides really valuable information. That's why PGO is a thing for C++.
I think that might also change as they add proposal and roadmap items like "direct access to the DOM"[1] to WASM.
Proposals: https://github.com/WebAssembly/proposals
Roadmap: https://webassembly.org/roadmap/
[1] https://github.com/WebAssembly/interface-types/blob/master/p...
Edit: Overall, the proposals seem to be pushing WASM closer to being a general purpose VM (VM as in JVM, not KVM).
The original essay shows exactly that, the one here shows pretty epic optimisations (and needs for an excellent knowledge of the platform) to get to "naive" WASM performances.
If that were true, nobody would be working on WASM because there would be no point.
"For a hot math loop doing floating point calcs (mandelbrot calc), however, I've seen javascript end up with identical performance compared to WASM. It's really pretty nuts. "
That's easy mode for a JIT. If it can't do that it's not even worth being called a JIT. That's not a criticism or a snark, it's a description of the problem space.
The problem is people then assume that "tight loops of numeric computation" performance can be translated to "general purpose computing performance", and it can't. I have not seen performance numbers that suggest that Javascript on general computation is at anything like C speed or anything similar. I see performance being more rule-of-thumb'd at 10x slower than C or comparable compiled languages. Now, that's pretty good for a dynamic scripting language, which with a naive-but-optimized interpreter tends to clock in at around 40-50x slower than C. The competition from the other JITs for dynamic scripting languages mostly haven't done as well (with the constant exception of LuaJIT). But JIT performance across the board, including Javascript, seems to have plateaued at much, much less than "same as compiled" performance and I see no reason to believe that's going to change.
Wasm has much more potential than just optimization. It opens up capabilities like:
* Using other languages besides Javascript, or languages that just compile to JS
* Passing programs around vs passing data around
* Providing a VM isolate abstraction that's embedable in your existing language/ runtime
This might be the most under appreciated comment I have seen all day. While Typescript did some wrangling to bring a bit a bit of "normalcy" to Javascript, the consistency that C++/Rust bring from a coding language point is (for me) enough to have gottem me interested in the use of WASM. Maybe I'm just lazy or have a strong dislike of the headache I get from learning JS compared to switching over to other languages.
We need to redo all this work for WASM precisely because there is no such thing in general as "as fast as C Javascript". Javascript itself is not a suitable target for all those things. A 10x performance haircut off the top, in addition to what it costs to get to even performance that good (JITs tend to eat RAM like candy to get their performance increases on dynamic languages), just wasn't an acceptable base for the things you're talking about.
It’s also new. Maybe there is potential here for WASM to get faster.
WASM will probably get faster, but I don't necessarily expect Rust-generated WASM or other static language-generated WASM to get that much faster. I think most of it will come from JIT'ing stuff that will mostly affect dynamic language-generated WASM. That said I do think Javascript's JIT is going to be fairly close to the limit of what you can see in WASM either, because it is implausible that you could "just" compile the JS to WASM then use the WASM optimizer to do better than the Javascript JIT. The Javascript JIT has everything the WASM optimizer would have, and more, since there would inevitably be loss in the translation.
WASM may not significantly outperform the JIT of one particular browser on a given scenario but you are more likely to get homogeneous performance across different browsers.
We did that! Yea chrome basically keeps right up, of course not accounting for SIMD
(anyone know why they started with narrow SIMD? surely it's easier to emulate 512-bit simd with 128-bit simd instructions than try to compile 128-bit simd to take advantage of 512-bit simd?)
All that detective work? Is UNDOING what javascript do, idiomatically!
Can be argued that JS is "fast" after all that, but instead, that show Rust give it for free, because is idiomatic there.
The problem with JS and other mis-designed languages is that do easy things are easy, but add complexity is easier. And the language NOT HELP if you wanna reduce it.
In contrast, in Rust (more thanks to ML-like type system), the code YELL the more complex it becomes. And Rust help to simplify things.
That is IMHO the major advantage of a static type system, enriched on the ML family: The types are not there, per-se, as performance trick, but for guide how model the app, then you can model performance!
P.D: In other words? If JS were alike rust (ala: ML, plus more use of vectors!) it will be far easier to a) not get on a bad performance spot in first place b) easy to simplify things with type modeling!
It would be better if we had JS and another language that were both usable in browsers to interact with the DOM, where this other language was more low-level. And this was the original plan — the “other language” was originally going to be Java. (Java applets were originally capable of manipulating the DOM, through the same API surface as JavaScript!) Then people stopped liking Java applets, so it moved to thinking about “pluggable” language runtimes, loadable through either NPAPI (Netscape/Firefox) or ActiveX (IE), enabling <script type=“foo”> for arbitrary values of foo. This effort died too, both because those plugin systems are security nightmares (PPAPI came too late), and because browsers just didn’t seem willing to standardize on a plugin ecosystem in a way that would allow websites to declare a plugin once that enables the same functionality on all past, present, and future browsers, the way JS does.
Eventually, we acknowledged that all browsers had already implemented JavaScript engines (incl. legacy browsers on feature-phones) and so it would be basically impossible to achieve the same reach with a latecomer. So we switched to the strategy of making browsers’ JavaScript engines work well when you use in-band signalling to (effectively) program them in another language; and we called the result WASM.
This isn’t the cleanest strategy. What’s great about it, though, is that (the text format of) WASM will load and run in any browser that supports JavaScript itself.
In short? mis-designed. JS was fora use case, and have become more and more for other uses case (even backend!). This is also the problem with html, css, dom.
Plus, some WATs are part of rushing it, and others for mismatch in the uses cases.
I'd point to not making async calls synchronous-to-the-caller by default as another pretty bad design mistake. The way so many JS files start nearly every line (correctly! This isn't even counting mis-use of the feature, which is also widespread!) with `await` is evidence of this, and so's all the earlier thrashing before we had `await` to deal semi-sanely with this bad design decision.
The original scoping system was just bad. We have better tools to make it suck less now, but it was designed wrong originally.
IIRC, some very clever JITs do a waltz here:
1. replace the user’s “syscall” into the runtime, with emitted bytecode;
2. WPO the user’s bytecode together with the generated bytecode;
3. Find any pattern that still resembles the post-optimization form of part of the bytecode equivalent of the call into the runtime, and replace it back with a “syscall” to the runtime, for doing that partial step.
Maybe clearer with an example. Imagine there’s a platform-native function sha256(byte[]). The JIT would:
1. Replace the call to sha256(byte[]) with a bytecode loop over the buffer, plus a bunch of logic to actually do a sha256 to a fixed-size register;
2. WPO the code with that bytecode loop embedded;
3. Replace the core of the remnants of the sha256-to-a-fixed-size-register code, with a call to a platform-native sha256(register) function.
There's similarly a cost in crossing JS<>WASM boundary, so WASM doesn't help for speeding up small functions and can't make DOM-heavy code faster.
Native code is opaque to the JIT, which means no inlining through native code (inlining being the driver for lots of optimisations) and no specialisation. This means if you have a JIT native code is fine for "leaf" functions but not great for intermediate ones, as it hampers everything.
When ES6 compatibility was first released, the built-in array methods were orders of magnitude slower than hand-rolling pure JS versions.
This issue is one of the things Graal/Truffle attempts to fix, by moving the native code inside the JIT.
Then in 2019 I rewrote the whole library from scratch (again), this time working on a MacBook using Chrome for most of my develop/test stuff. The library is now much faster on Chrome! I assume what happened was that I was constantly looking for ways to improve the code for speed (in canvas-world, if everything is not completing in under 16ms then you must own the humiliation) which meant that prior to 2019 I was subconsciously optimising for FF; after 2019 for Chrome.
> it requires very detailed knowledge of the JS engine internals and may require giving up most high-level features that define the language
I can't claim that I have that knowledge. Most of my optimisations have been basic JS 101 stuff - object pools (eg for canvas, vectors, etc) to minimise the stuff sent to garbage, doing the bulk of the work in non-DOM canvases, minimising prototype chain lookups as much as possible, etc. When I tried to add web workers to the mix I ended up slowing everything down!
There is stuff in my library that I do want to port over to WASM, but my experiments in that direction so far have been less than successful - I need to learn a lot more WASM/WebAssembly/Rust before I make progress, but the learning is not as much fun as I had hoped it would be.
> Javascript can be made surprisingly fast
The speeds that can be achieved by browser JS engines still astonish me! For instance, shape and animate an image between two curved paths in real time in a 2D (not WebGL!) canvas:
- CodePen - https://codepen.io/kaliedarik/pen/ExyZKbY
- Demo with additional controls - https://scrawl-v8.rikweb.org.uk/demo/canvas-024.html
JS optimizations are implementation defined. They are also largely undocumented. You will have to read source code if you want to know the truth, or write microbenchmarks to probe the VM behavior.
Your optimization from today can be deoptimized tomorrow. There are no documented guarantes that certain code will remain fast.
Knowing about inline caching, shapes, smis, deoptimizations and trampolines... does helps. But the internals are unintuitive.
Instead, I can save myself that time and use WASM instead.
There are so many things that can trigger a deoptimization that I would rather ignore them all and do it in WASM instead.
Here's some benchmarks. But in summary, for small simple hot loops, chrome is just as fast using js, but for firefox wasm gives big performance improvements.
Of course, if you can use SIMD instructions, then wasm will win, but that's less fair.
Hyperscript is an interpreted programming language written on top of javascript with an insane runtime that resolves promises at the expression level to implement async transparency.
Head to the playground and select the "Drag" example:
https://hyperscript.org/playground/
And drag that div around. That whole thing is happening in an interpreted event-driven loop:
repeat until event pointerup from document
wait for pointermove(pageX, pageY) or
pointerup(pageX, pageY) from document
add { left: `${pageX - xoff}`, top: `${pageY - yoff}` }
end
The fact that the performance of this demo isn't absolutely CPU-meltingly bad is dramatic testament to how smoking fast javascript is.The fact that you're excited at how dragging a single div around actually works as expected is a testament to how horrifyingly low your expectations are.
Maybe you don't need Rust and WASM to speed up your JS - https://news.ycombinator.com/item?id=16413917 - Feb 2018 (181 comments)
Our conclusion was that wasm was consistent across browsers whereas js wasn't. Further, If you can use simd, wasm is faster. Also, v8 is way faster than spidermonkey.
And since you are talking about the fastest mandelbrot on the web, here is my contribution https://mandelbrot.ophir.dev
It renders in real time and is fully interactive, all in pure js.
The reason I said fastest is cause I wrote the hot loop with SIMD instructions. I'll be adding in period checking and stuff in a couple months time, will be sure to refer your work then, thanks.
Perhaps, but you do need it in order to have an excuse to use Rust and WASM. True story, this is what I did last weekend.
There are some game engines in Rust for simpler 3D games. I haven't tried them. 2D games seem to be well supported in Rust. I'm working in Rust, writing a client for a virtual world, and I have to say that the 3D graphics ecosystem for that kind of thing in Rust is not quite there yet. I'm using bleeding edge packages (Rend3->WGPU->Vulkan) where I'm in regular contact with the developers. I'm doing it this way mostly because I want to see if I can get better performance than the existing single-thread C++ client, which is compute-bound in the main thread.
Generally though, I think wasm represents an opportunity to break the strangle-hold JavaScript has had on the frontend ecosystem for nearly 3 decades. Typed, compiled languages typically have huge advantages in terms of safety and performance over dynamically typed interpreted languages. And now a lot of high-level features and syntactic sugar that used to be only available in dynamic languages is becoming available in systems languages.