We've been lied to: JavaScript is fast
jyelewis.com
jyelewis.com
JavaScript is probably like 10x or 100x slower than C or similar languages when you write something bigger and any of the following happens:
- the ratio of code size to run time is too big. Then the jit can’t keep up.
- you use a lot of value types. Them be structs in C, no need for allocation. In JS them be objects. The JS VM will try to escape analyze them, and it will succeed some of the time, but fail enough of the time to cause massive punishment.
- you churn objects while having a large base heap size and the generational hypothesis holds only a bit. Then you’ll wait for the GC a lot.
- you have a class hierarchy with many descendants and you often access properties or call methods on the base type. Then vtables or whatever work great but JS inline caches blow up.
- probably lots of other conditions.
Write enough code and at least one of these will hold and you’ll be slow in JS, fast in C.
(Source: I work on JSC and I implemented a lot of its optimizations. Benchmarks like the ones in this post are the sort of thing my JITs eat for breakfast. It’s cool to see people throwing me softballs but I like to be honest about what the technology I work on is capable of.)
JavaScript, Java, C#, generally all JITed languages will eventually end up exhausting the size of their JIT code caches. You just gotta hope that your hot paths are contained in there and that the JIT balances its hunger for deep inlining with the need to not overcompile. Because there's nothing worse for a JIT when the entire application is just a tepid soup and there are no hotspots. Then you aren't getting escape analysis, you end up with tons of polymorphism, and generally, performance suffers.
Despite 20 years working on JITs and dynamic optimization, I am more convinced than ever you just can't beat static compilation and programs designed to not over-allocate, to not overabstract with too much polymorphism, closures, and heavy allocation. Programs still need to be a bit miserly to get maximum performance.
But the thing is that Java is pretty close to native. Primitive types, object allocation, threads etc. all correlate almost directly to how CPUs and OSs work. So you get a sense of performance and what's expensive/isn't.
With JavaScript there are so many complex abstractions and constructs it's really hard to get a sense of what the final ASM will look like in an optimal case. Unfortunately, some Java changes (e.g. Valhalla) aim to bring some of that lack of clarity into Java. Still, thanks to typing and a simpler syntax Java is MUCH better positioned here than JavaScript.
I'd like to add that this argument exists for pretty much all "slow" languages and that "if you're careful" often can mean "throw away all upsides your language provides" [1].
Personally, I find these discussions tiring because the ones arguing such points are usually the kind of person that doggedly clings to the "wrong tool for the job".
For example: Although I see rust as the better alternative in many places, I personally find C++ to still be a better language than Java. That doesn't mean I'll insist that Java has no worth or that every program should be written in lower level languages. Horses for courses.
Why can't we just go the Python way and accept the shortcomings of the languages we hold dear?
[1]: Edit: For GCed languages, this often is "avoid the GC at all costs and essentially write less readable C".
That doesn’t solve the problem of course. I have this metric that I use to philosophize about this: SIPS, or static instructions per second. If your program has low SIPS (I.e. it’s the benchmark of HotSpot’s wet dreams or OP’s post) it means that the total code is small and it runs a long time - so the JIT will go on a rampage early on and then you’re good. If the program has high SIPS, then the JIT will still be catching up when the program terminates. Most real things that aren’t web servers have low SIPS.
I haven't really paid attention to V8's policies for years, but generally speaking, it still does this to some extent.
I think of SIPS as basically the resident set size for code. Depending on how the system ages your code, you might be stuck with unused JIT code until it can be recycled, if at all. V8 used to have code aging and the implementation was a huge bug farm. I am not sure if that survived the switch to TurboFan/Ignition. Of course, V8 GC's code about as aggressively as it compacts the old generation.
WebAssembly also introduces its own issues since it’s a BYORT system (bring your own runtime). So, there’s more to JIT (your language’s whole runtime) and more to hold in memory (your language’s whole runtime). You might say, “but pizlonator, every language has a runtime”, to which I’d say: yes but usually that shit gets shared by every running instance somehow. In WebAssembly every instance pays for its runtime’s memory footprint and it’s quite likely that every instance has a different runtime so the code isn’t shared either. Basically, WebAssembly is knee-capped on memory footprint by design. JS isn’t.
WebAssembly has no truck with JIT. Just-In-Time compilation is all about not performing optimisation ahead of time, but only doing it when the code in question is used, and guiding the optimisation by how the code is being used. In WebAssembly, the conversion from the binary .wasm format to machine code is a comparatively lightweight process with no real optimisation of this form: rather, such optimisations must be done as part of producing that .wasm blob (in regular compilation possibly with profile-guided optimisation to go even further than JIT can).
So instead, when you’re using something like Rust (as distinct from, say, compiling CPython or V8 to WASM), what you’ve got is a fairly small amount of what you’d probably call runtime code (allocator, panic mechanism, str, Vec<_>, some other parts of the std crate), probably something like 25KB (or with a little care and compromise, more like 5KB) except when Unicode tables are required, compiled from WASM byte code to similarly-sized (though I don’t know the real ratio) machine code faster than it can be downloaded. That’s code memory usage; for the data memory usage, well, your Rust/WASM will normally blow your JavaScript out of the water there with much more efficient packing of data into memory, even if you’ve got a fair bit of overallocation.
The fact of the matter is that the runtime parts which can be shared for JavaScript are actually not all that large, and routinely dwarfed by included libraries (React, &c.). WebAssembly is by no means knee-capped on memory footprint; rather, so long as you’re using it in a sane way (Rust, standard WASM optimisation and code-shrinking techniques, that kind of thing), there’s a pebble in the way that you’ll notice if you’re a beetle, but if you’re even rabbit-sized you probably won’t even notice it.
As regards pbadenski’s comment: WebAssembly does not run in the JavaScript virtual machine, it’s a completely separate thing that is merely able to call and be called from JavaScript via a foreign function interface. The backing memory buffer may also be accessed from JavaScript, but that’s immaterial in the consideration.
Compiling wasm to native is only faster than downloading if you compile without optimization. That’s common in wasm VMs but then there’s an optimizing JIT that runs adaptively later, just like a JS or Java VM would do.
The fact that JS VMs share the ICU implementation between instances is a huge deal. That’s not the only thing that gets shared. Also, it’s not about just sharing code; it’s about sharing memory for the runtime’s state and for allowing elastic reuse of space for objects. In wasm the sharing is page granularity at best.
I can confidently say that games and apps in the browser and Electron are noticeably slower than other apps. You don’t see many web games because the graphics required for games today can’t really be rendered on the web in 20+ FPS, at least until WebGL2/WASM/WebGPU get better support.
But JavaScript is fast enough. Because the vast majority of programs, especially websites, don’t need really fast code like C. They don’t need to redraw every frame, don’t need 3D capabilities, don’t need to perform expensive computations, etc.. If you need a fast program that does those things 99.9% of the time you can just make it a real app, or you can write the fast parts in WebAssembly.
At the end of the day computers are very fast, so a program could be written in any language as long as it isn’t doing anything super performance-needing and the compiler/interpreter has has half-decent optimization.
Of course, in current gen browsers and on current gen hardware, sunspider is so fast it's more noise than measurement.
[1] https://www.zdnet.com/a/img/resize/3062e92bf2f9f018f49723939... Full size image url is broken :P
This helps everywhere: mobile (better battery life, smarter apps), desktop (more apps open without swapping), server (more compute in the same footprint), and embedded (smarter devices, less power, cheaper hardware).
Oh, and the Web API has has stuff you wouldn't expect like USB, Bluetooth, MIDI, vibration, TPM etc.
Here's a Web Bluetooth example: https://googlechrome.github.io/samples/web-bluetooth/device-...
In my experience the great majority of perceptible slowness in browser apps comes from DOM reflows, not JavaScript
A few years ago I noticed most websites just seemed slower than they should on my computer. The culprit turned out to be metamask (an etherium wallet extension). It was adding 700 kilobytes of javascript to each page load thanks to web3. It also exposed my etherium address to any website that asked.
The JS was cached, but just parsing that much code added noticeable lag to every page load. Everything felt zippy once I uninstalled that.
Certainly the web as a platform is slower than a precisely-built thing honed for speed for a particular use case, but speaking generally the web is only perceptibly slow when you write bad code (or are waiting for the network). Which is admittedly rather common.
Nonsense. There's nothing wrong with the graphics you can run on the web. Web games flourished before the iPhone. And it seems safe to assume that they died because of the iPhone.
It was a horrific development in the space, since mobile games differ from flash games in being (1) more expensive and (2) worse.
Well, my problem with the testing process that goes into many programs is that they only test with decent hardware. What about those who do not have the latest, state of the art computer or smartphone?!
I notice that the author doesn't give any links to claims that JavaScript is slow, and the first two pages of search results are either discussions of how fast JS is or guides to JS performance. Who exactly was lying?
I should also add a massive asterisk to "never slow", clarifying that it's usually within a couple of orders of magnitude of compiled code, and faster than that for glue language tasks that let it call native code for the heavy lifting.
Having written a lot of code in both Python and JavaScript I disagree here. It's not really the GIL slows down Python compared to JavaScript, that's more down to the lack of a comparable JIT. Yes there's PyPy but it has received a fraction of the investment of V8 and Python is a more complex language to create a good JIT for. Support for multithreading (albeit without parallelism) may be one of a number of factors those complicating factors.
JavaScript doesn't need a GIL because it doesn't support multithreading. WebWorkers are more akin to Python's multiprocessing since objects are not shared between threads.
You have to be careful doing anything compute intensive (e.g. JSON.parse of a large blob of JSON) when using cooperative (async) rather than preemptive (threads) concurrency since you may inadvertently block the event loop, substantially increasing latency for other tasks. Handing off such tasks to a thread pool is harder in JS than Python since you can't easily deferToThread as you can in an async Python program.
Cooperative multitasking has its advantages too of course. An event loop provides a clear boundary for efficiently batching calls which can be awkward with synchronous code. Folks have been writing network servers and in Python using asynchronous techniques for decades now. Twisted is almost 20 years at this point and asyncore was added to the standard library even earlier.
`findSmallestPositiveValue([2,11])` gives the wrong value for the supposedly better function.
> Even if the second function takes twice as long as the first, we are in the realm of nanoseconds.
How can you possibly know this? You can make either function take arbitrarily long by increasing the size of the input. The second one scales worse. I see no reason to assume an upper bound on the input size given. If you want the behavior to be obvious at a glance, just name the function what it does. Just like it already is. Or leave a block comment.
The argument here is that javascript is so fast that it doesn't matter what you write because you don't have to worry about that. I just don't know how to reconcile that with the fact that I regularly see websites that have noticeably slow javascript "startup" times.
I think this is the point, the second one scales worse - absolutely (ignoring the bug you mentioned) however it really is more readable, and because V8 is so fast, why not use the second version?
Block comments and function names are important, but if someone needs to modify the code to add or change functionality nothing beats simple & readable code.
In some cases, there are persuasive arguments to favor readability over performance, but I don't find this particular example very convincing.
True, same as any language or applications fast CPUs shouldn't be a free chance to completely ignore writing efficient code. Although I bet those websites are slow because they are doing silly things like adding 1000 items to the DOM one by one rather than being slow because someone didn't optimize their use of .forEach()
Javascript is very fast, but what is not fast is the frameworks that use it or the ways people use those frameworks. If people with similarly minded thinking as those who develop desktop only apps with native languages, would write javascript only apps, I would say the performance would be very close.
But many developers have no understanding of how to write performant code, as they have used to just using some heavy frameworks that handle all that stuff for you, so easily you become limited by that thinking.
That is at least my view into this world, where naive web-developers who have like 2 years of software development can get to deploy stuff to production. Iteration speeds need to be fast, so a lot more people are employed without the skills to really understand what is happening, thus resulting in poorly executed services. Or maybe it's just not a priority.
Also, many technically inclined people just don't get it, that what is important is that you can execute a function, the speed is not really so important in the end, even though we would like it to be. We tend to live in a bubble, those who have dwelled deeper into the operating ways of computers, that we think that everybody else is like that too, or should be.
Many are just doing their jobs, and that does not include learning how a CPU or memory works, but it might be limited to learning how a Vue.js framework works or how to use React.
Doesn't that just mean it can take a long time for the browser-application to download all of its scripts?
That can certainly take a long time but it depends on 1) The speed of the web-server(s) serving those scripts 2) The speed of the network you are connected to
So web-sites taking a long time to execute their JavaScript on your PC in the browser does not say much about the speed of JavaScript the language.
But it is a practical concern of course. Which brings to my mind the fact that JavaScript programs can exhibit extreme parallelism, because they can execute in millions of browsers at the same time, think SETI-search, or unauthorized crypto-mining.
So in practice JavaScript can have a huge computational throughput, in other words great speed, because it can execute on multiple clients at the same time, easily.
I remember node.js benchmarks in the front page of hn showing it to be as fast as hotspot jvm. And this was pre 2012.
And since node is single threaded, you had to run n.cores instances of your app.
For the longest time as a kid what I was lied to was that Java was much slower than Ruby/Python and my perception was painted by the startup times.
I don't wish to start a language flame war here. We're beyond that.
But by what measure? Doesn't this depend on what you're doing?
By using JavaScript you are throwing away something like 2-3x.
There will be cases where JS can keep up because it reduces to a local loop which gets optimized the same way as any other language, but all those little other bits add up.
Same with Python. If you do everything in numpy you can be very fast, but what happens if you don't? You could use a fancy jit but now you have a deeper stack with no control over the emitted code - nice if you already have the code, but not ideal.
All languages have their pain points, and modern JS/TS, while undisputably having their quirks, aren't particularly more quirky than JS (IMHO)
There certainly are stable packages, but sometimes to do what you need to do there is no other option.
As a result, I'm convinced that software written in JS is much more likely to actually perform long-running (IO-bound) tasks in the background/in parallel, simply because it isn't a huge pain to do it.
The actual execution speed matters a lot less at that point.
Profiling NodeJS is futile because of its async nature. You end up with a lot of noise and no substance. I'm looking forward to project Loom which would bring Java threads into hybrid green/native mode. That would deliver throughput as fast as NodeJS but with the performance of Java and clarity of simpler stack traces.
For server applications handling many requests in parallel anyways, it's a different story.
This brings to my mind the current logistics crisis in the global trade. Lots of parallel tasks and traffic jams and all of a sudden stores don't have stuff and prices are going up and we don't know when will it be back to normal. That is what Node.js can feel like, because it is hard to understand what tasks execute when and who is waiting for what.
function main() {
let myNum = 0;
for (let i = 0; i <= ITERATIONS; i++) {
myNum *= i;
myNum++;
}
}
I'm a little surprised that turbofan doesn't JIT that into an empty function, given that it has no return value or side effects. https://brython.info/static_tutorial/en/index.html
https://github.com/qquick/Transcrypthttps://github.com/sagemathinc/JSage/tree/main/packages/jpyt...
I am curious why this did not happen or what optimization flags were passed in.
If it was prevented optimization through flags, is this a fair test?
Then I added a printf statement and the algorithm came out without too many strangeness.
When I reduce the amount of iterations, the compiler inserts a constant where the calculation would've taken place. The amount of iterations in the benchmark seems to be too high for the compiler to evaluate and optimize out. The clang seems to stop any full evaluation at exactly 101 iterations, meaning there's probably a default limit somewhere puts the limit on a nice, round 100.
I've taken the code, increased the amount of loops by a factor of 10 (to make the differences more pronounced), added a printf/console.log to make sure the loops themselves don't get thrown away (clang does that with -O2) and on my laptop (i7-10750H) the C code, compiled with clang 12, runs 10000000000 iterations in 9.462s whereas the same number of iterations in Node 16 runs in 19.990s. That's more than a 2x execution duration, with more than a 100% speed difference.
For comparison: Java runs in about 9.846s, from the command line (java code.java, no compilation step). C# (dotnet 5) runs in about 9.736s after compilation; compilation takes about a second. Rust runs in about 9.435s after about 0.47s of compilation. Kotlin runs in 9.640s, but it took a few seconds to be compiled into a JAR first. Python 3.9.7 takes forever, but I think that's because its arbitrary length number implementation is trying to make it output the correct result instead of faking it like the other programs are doing. PHP 8 also didn't really stand a chance at over 90 seconds.
JS would probably win in a huge code base consisting of mostly dead code (node_modules, anyone?) where compilation would get in the way of quick edit-run-verify loops, but only starting from scratch without a compiler cache set up. From what I can tell, JS certainly isn't _slow_ like PHP, but it isn't _fast_ either. It's somewhere in between the old interpreters of old and compiled/runtime-JIT'ed languages.
If anything, this algorithm benchmark would indicate that you're probably better off running C#/Java rather than C/Rust because of the negligible performance difference with the huge benefits of language safety with no effort. This benchmark is far from normal program code, of course, so it doesn't really prove anything.
It should be noted that NodeJS outputs the result as "Infinity" whereas C and other languages creates a value that's clearly been bit-wrapped and overflown quite a bit. This can be an advantage (because Infinity + anything = Infinity) or a disadvantage (float math) but that's just how the language works.
Edit: looking at the rest of the code, I've also benchmarked the loop vs functional approach (Node 16, same device).
// Generate numbers
numbers = [];
for (let i = -10000000; i < 10000000; i++) numbers.push(i);
lowest = numbers.filter(x => x > 0).sort((a,b) =>a-b)[0];
console.log(lowest);
^ this runs in 1.007s numbers = [];
for (let i = -10000000; i < 10000000; i++) numbers.push(i);
let lowest = Infinity;
for (const i of numbers) {
if (i > 0 && i < lowest) {
lowest = i;
}
}
console.log(lowest);
^ this runs in about 0.629sI wouldn't call a near 40% speed difference "only 2ns". The functional approach is very comfortable to program in, but it comes at a real cost and should definitely be avoided complex in algorithms.
Java handles streams very poorly. However, LINQ is quite fast:
var numbers = new long[20000000];
for (var i = -10000000L; i < 10000000L; i++) numbers[i + 10000000L] = i;
Console.WriteLine("{0}",
numbers.Where(x => x > 0).OrderBy(i => i).First()
);
^ this runs in 0.346 seconds.That doesn't make it quite optimal, though:
var numbers = new long[20000000];
for (var i = -10000000L; i < 10000000L; i++) numbers[i + 10000000L] = i;
long lowest = 9999999999L;
foreach (var i in numbers) {
if (i > 0 && i < lowest)
lowest = i;
}
Console.WriteLine("{0}", lowest);
^ This runs in 0.151sThese examples aren't very conclusive either. Rust's loop version runs in 0.100s and a shitty iterator version that collects the entire iterator and sorts it runs in 0.130ms. Again, the real fight here seems to be between C# and something closer to the metal.
$numbers = [];
for ($i = -10000000; $i < 10000000; $i++) $numbers[] = $i;
$lowest = PHP_INT_MAX;
foreach ($numbers as $i) {
if ($i > 0 && $i < $lowest) {
$lowest = $i;
}
}
1.08 secondsPHP Optimized
$lowest = PHP_INT_MAX;
foreach (range(-10000000, 10000000) as $i) {
if ($i > 0 && $i < $lowest) {
$lowest = $i;
}
}
0.56 secondsMake sure to enable JIT (tracing)
Node for loop version for me took 1.25 seconds (I have a very old rig)
php --version
PHP 8.0.11 (cli) (built: Sep 21 2021 18:25:57) ( NTS Visual C++ 2019 x64 )
Copyright (c) The PHP Group
Zend Engine v4.0.11, Copyright (c) Zend Technologies
with Zend OPcache v8.0.11, Copyright (c), by Zend Technologies
node --version v16.1.0
Note: I measured time within the script and not outside to avoid any bootup time.Also, why ignore memory usage and JS VM startup time? Why ignore the massive dependency tree that comes with JS frameworks? There's multiple factors to consider when assessing language performance. JS is far less efficient than other languages...but that's honestly fine. There are plenty of use cases for the language where your application doesn't have to be incredibly performant.
I think desktop GUI apps also have browser rendering engines chewing resources.
Followed by a search smallest function which is linear in time optimized to a function using filter/sort which is n*log(n) in time. Impressive indeed.
> int main()
The easiest way to find someone who's inexperienced in C is to find someone who declares a function that takes no arguments with an empty pair of parentheses. In C but not C++ you need to write "void" inside the parentheses.
> myNum *= i;
Signed integer overflow is undefined behavior. You are lucky if the compiler didn't just replace the whole thing with __builtin_unreachable().
Next the function doesn't use the computed variable. A compiler can just optimize the whole thing out.
The point I'm trying to make is that C is a difficult language to write; especially so when you want to benchmark.
Compilers are quite chaotic when you get deep in the passes.
C11§6.7.6.3p14:
> An identifier list declares only the identifiers of the parameters of the function. An empty list in a function declarator that is part of a definition of that function specifies that the function has no parameters. The empty list in a function declarator that is not part of a definition of that function specifies that no information about the number or types of the parameters is supplied.
That declarator is part of a definition. Even if it were not, the elision of a function's formal parameter list does not matter very much if, like 'main', that function is usually not explicitly called.
Are there languages that are better at doing IO than others?