All optimizations in the post can mostly be divided into three large groups:
1) algorithmic improvements; 2) workarounds for implementation independent, but potentially language dependent issues; 3) workarounds for V8 specific issues;
You need to think about algorithms no matter which language you write in, so we don't need to talk much about the first group. In the post it is represented by sorting improvements (sorting subsequences rather than the whole array) and by discussions of caching benefits (or lack of them there-off).
The second group is represented by a monomorphisation trick: the fact that the performance suffers due to polymorphism is not really a V8 specific issue and it is not even JS specific issue. You can apply this approach across implementations and even languages. Some languages apply it in some form for you under the hood.
The last group is represented by argument adaptation stuff.
Finally an optimization I did to mappings representation (using typed array instead of an object) is an optimization that spans all three groups. It's about understanding limitations and costs of a GCed system as whole.
Now... Why did I choose the title? That's because I think group #3 represents the issue that should and would be mostly fixed over time. While groups 1 and 2 represent universal knowledge that spans across implementations and languages.
Obviously it is up to each developer and each team to choose between spending N rigorous hours profiling and reading and thinking about their JavaScript code, or to spend N hours rewriting their stuff in a language X. What I want is:
a) that everybody was fully aware that the choice even exists; b) language designers and implementors worked together on making this choice less and less obvious - which means working on language features and tools and reducing the need in group #3 optimizations.
I suspect some respondents will say, or believe, that this falls under the category of "you have to know the VM/runtime intimately to get good performance".
I don't think that's true. If you know the general problems with polymorphic call sites, you can (in JS and many other languages) check to see if the runtime is applying optimizations of this kind by being explicit. If it helps, you can get a free optimization by setting explicit arity/dispatch just because you happened to know about the issues surrounding polymorphic dispatch. That's a case of fundamental knowledge speeding up/improving your progress; not a case of "you need to know the VM guts like the back of your hand to make code fast".
Here's a crazy thing I recently learned: apparently monomorphism isn't just "object with identical keys", apparently (at least in Chrome), the order in which you declare those keys matters. According to this presentation from 2015[0], adjusting the following lines in the Octane/Splaytree benchmark so that node.left and node.right are always assigned in the same order resulted in 15% better performance:
var node = new SplayTree.Node(key, value);
if (key > this.root_.key) {
node.left = this.root_;
node.right = this.root_.right;
...
} else {
node.right = this.root_;
node.left = this.root_.left;
...
}
Now, I assume that this out-of-order thing was actually done on purpose, to benchmark how the JIT handles code like this. Further evidence for that is that the SplayTree constructor[1] does not feature a left and right key either: SplayTree.Node = function(key, value) {
this.key = key;
this.value = value;
};
Still, I wouldn't be surprised if it was common for real-life code to accidentally have objects that should have the same hidden class end up with different ones because of this.[0] http://mp.binaervarianz.de/fse2015_slides.pdf
[1] https://github.com/chromium/octane/blob/master/splay.js#L390
Although I suppose you're already required to have two separate hidden classes to distinguish these two kinds of objects anyway.
> with a bizarre exception for arrays
Wow, you weren't joking with how bizarre this gets:
var a = {};
var b = {};
a.a = 0;
a.b = 1;
a[0] = 0;
a[1] = 1;
b.b = 1;
b.a = 0;
b[1] = 1;
b[0] = 0;
Object.keys(a); // Array [ "0", "1", "a", "b" ]
Object.keys(b); // Array [ "0", "1", "b", "a" ]
PS: Thanks for making a great shell :)Chrome/V8's team found they could get substantial performance improvements by diverging from the de facto standard, without too much of a cost to web compatibility.
See: https://stackoverflow.com/questions/5525795/does-javascript-...
The order is only sometimes guaranteed, of course, because JavaScript. (But critically, for this discussion, it is important that it is sometimes guaranteed, because it forces that information to be stored by the VM.)
// the keys of an object
export function keysOf(obj) {
let keys = Object.keys(obj);
keys.sort();
return oneOf(keys);
}
Context: the rest of the code takes an object representing a schema, and creates two functions. One that can turn any object fitting that schema into an array, with positions indicating which key they originally belonged to, and another one that can reverse the process.Could this also explain why I have no consistent order of properties when viewing state with the Redux dev-tools? Instead of Object.assign I use my own simplified merge code[1]. Maybe if I also make that use a sorted set of keys, the devtools will become more consistent in their presentation (and it might result in more consistent hidden classes too).
[0] https://github.com/linnarsson-lab/loom-viewer/blob/master/cl...
[1] https://github.com/linnarsson-lab/loom-viewer/blob/master/cl...
The key takeaway of your commentary on asm.js five years ago (http://mrale.ph/blog/2013/03/28/why-asmjs-bothers-me.html) is the same as the key takeaway from this post: we haven't reached the end of JS engine performance improvements, and if we apply the same rigor to JS development that we apply to C++ or C or Rust or some other language the results are definitely surprising!
https://github.com/WebAssembly/host-bindings/issues/11#issue...
If you think the assembly in this article is bad, WASM/host-bindings appear to be even worse.
The point of rust+wasm benchmarks is that one can write reasonable, maintainable, functionality-focused code (not to mention all the rust-specific benefits) and get good performance out of the box.
Radix sort for instance comes to mind.
This isn't unique to this situation either. If you're writing Python and your choice is to either deep dive into various hacks to maximize performance of your pure python or learn Rust/C/C++ and call out using an FFI mechanism, if you have the time to space, I think it's almost always better to learn the new language. There are many benefits to learning new languages beyond just the different performance characteristics, so if you can afford to take that path, I think it's usually a good choice to do so.
As far as I know Rust has just two major success stories right now: Firefox 57 (Servo) ripgrep (even included in Microsoft Visual Studio)
I wouldn't say this is entirely wrong, but proper cross-platform package and test management with Cargo is a reason alone that utilizing existing code is waaay easier in Rust. C++ has more existing code though, so they might hit your niche needs better.
In what concerns Windows development, NuGET and vcpkg are already a big improvement.
C++ had the advantage of being immediately adopted by the OS vendors for GUI development, although nowadays, with exception of Windows, its role has changed into just addressing the GPU.
So maybe one day we will get something like shaders, CoreGraphics, DirectX, SurfaceFlinger in Rust, but it will still need a couple of years.
Webrender is already something into this direction.
Here's a bunch more. Dunno what your definition of "major" is though: https://www.rust-lang.org/en-US/friends.html
I think people think too lightly of rewrites. Actually, I'd even argue that in some cases, rewrites are done because they're easier than actually fixing the problem. Yes, Rust and other languages will give you better performance out of the box, but at what cost?
I agree that rewrites are often taken too lightly, but if they address the original problems I think it would be more accurate to say that they are often needlessly expensive ways to solve problems that can also be solved in other, cheaper ways.
I'd also point out that in the case of many open source projects, finding an optimization consultant is not even remotely an option. For many of those projects, if performance is suffering, someone needs to step up and figure something out. Then the question becomes which approach can be applied by some contributor who's actually willing to do it. If you don't have someone who understands polymorphism in VM runtimes, I think in many cases you'd be well served by sprinkling some wasm on the problem. Of course this doesn't apply in all cases.
Yes. I wrote about this in some detail 2 years ago, Jitterdämmerung: