Show HN: Fast.js – faster reimplementations of native JavaScript functions
github.com
github.com
>I'd expect native browser methods to be an order of magnitude faster.
Built-ins need to adhere to ridiculous semantic complexity which only gets worse as more features get added into the language. The spec is ruthless in that it doesn't leave any case as "undefined behavior" - what happens when you use splice on an array that has an indexed getter that calls Object.observe on the array while the splice is looping?
If you implemented your own splice, then you probably wouldn't even think of supporting holed arrays, observable arrays, arrays with funky setters/getters and so on. Your splice would not behave well in these cases but that's ok because you can just document that. Additionally, since you pretty much never need the return value of splice, you can just not return anything instead of allocating a wasted array every time (you could also make this controllable from a parameter if needed).
Already with the above, you could probably reach the perf of "native splice" without even considering the fact that user code is actually optimized and compiled into native code. And when you consider that, you are way past any "native" except when it comes to special snowflakes like charCodeAt, the math functions and such.
Thirdly, built-ins are not magic that avoid doing any work, they are normal code implemented by humans. This is biggest reason for the perf difference specifically in the promise case - bluebird is extremely carefully optimized and tuned to V8 optimizing compiler expectations whereas the V8 promise implementation[2] is looking like it's directly translated from the spec pseudo-code, as in there is no optimization effort at all.
[1]: https://github.com/angular/angular.js/issues/6697#issuecomme...
That kind of optimization seems easily done with a builtin as well: have an option to return the array, and only do so if the JavaScript engine indicates that the calling code actually uses the return value.
Whereas the builtin ones implemented in C++ do keep all the code in there.
This is why even full implementations written in JavaScript can be faster than C++.
So not just JS - it seems language built-ins are often slowed down by bureaucratic spec compliance, and hand-rolling code can help you get a speedup.
I ran a profile. The code ran for 2.5 seconds, 4 ms of which in the micro-optimized js code, the rest updating the dom, five times all over again. Needless to say that i threw out all the micro-optimizations, halving the number of lines, and fixed it so the dom was updated once.
Anyway, the point i'm making is this: you should micro-optimize for readability and robustness, not performance, unless profiling shows it's worth it. I haven't known a case where pure (non-dom, non-xhr) js code needed micro-optimization for performance in half a decade.
[UPDATE: also I was reading SVG/DOM as an or]
For example, with IE on Win8, your DOM UI is actually composed of DirectComposition layers, and touch manipulations (pan, zoom, swipe, etc) are performed entirely on a separate thread from your UI+JS using DirectManipulation and a kernel feature called delegate input threads, controlling animations of the composition layers on yet another (composition) thread. You'll have a heck of a time matching that with canvas and mouse/touch/pointer events on your lone JS UI thread and the dozens of responsibilities it's signed up for.
Again, it depends a great deal on what you're doing. And it's a trade-off for sure... DOM is expensive, and if you're doing something too complicated then having panning and hit-testing on another thread doesn't help if the user ends up seeing large unpainted regions for a while, or if you make your startup time painful by trying to build a complex scene upfront. But you asked, so I answered :-)
There are also other important considerations:
1) It's easier (at least until we get some awesome UI frameworks on top of canvas, if that works out well).
2) It's compatible with older browser versions.
3) DOM layouts and controls support various accessibility features, input modalities, etc.
4) There's a huge library of tools/utilities/reusable components all designed specifically for it.
(I used IE/Win8 as an example because I'm familiar with how that works in great detail. Last I knew no other platforms were that aggressively optimized in those respects but presumably some are part way there)
Safari counts and profiles ALL function invocation, so if you go looking for hotspots with it's profiler, you always get the parts where many functions are called. I consider this harmful.
I stumbled over this. Maybe other developers did too.
[0] https://developer.mozilla.org/en-US/docs/Web/JavaScript/Type...
(edit): Yeah, the abstraction overhead is ridiculous. Here's the forEach() benchmark again, compared to an explicit for loop (no function calls):
// new benchmark in bench/for-each.js
exports['explicit iteration'] = function() {
acc = 0;
for (var j=0; j<input.length; ++j) {
acc += input[j];
}
}
Native .forEach() vs fast.forEach() vs explicit iteration
✓ Array::forEach() x 2,101,860 ops/sec ±1.50% (79 runs sampled)
✓ fast.forEach() x 5,433,935 ops/sec ±1.12% (90 runs sampled)
✓ explicit iteration x 28,714,606 ops/sec ±1.44% (87 runs sampled)
Winner is: explicit iteration (1266.15% faster)
(I ran this on Node "v0.11.14-pre", fresh from github).This is very often overlooked but extremely useful for implementations of fast algorithms in JavaScript that should scale to a lot of input data.
The next step for fast.js are some sweet.js macros which will make writing for loops a bit nicer, because it's pretty painful to write this every time you want to iterate over an object:
var keys = Object.keys(obj),
length = keys.length,
key, i;
for (i = 0; i < length; i++) {
key = keys[i];
// ...
}
I'd rather write: every key of obj {
// ...
}
and have that expanded at compile time.Additionally there are some cases where you must use inline for loops (such as when slicing arguments objects, see https://github.com/petkaantonov/bluebird/wiki/Optimization-k...) and a function call is not possible. These can also be addressed with sweet.js macros.
They have a structure like image[pixel].color
What I found was, always traversing the object structure was much slower than putting every pixel color as an argument in a function, that gets called every iteration.
So I had the impression, reasonable simple function, like a greyscale filter, get inlined by engines like V8.
for (key in object)
is for getting keys, in Firefox, meanwhile 'of' for (value of iterator)
is for getting values out of an Iterator.How do you mean when you state that for(of) is incorrect?
Perhaps this 'example' clears something up?
« for (value in {a:1,b:2,c:3,d:4,e:5}) console.log(value)
» undefined
"a"
"b"
"c"
"d"
"e"
« for (value of {a:1,b:2,c:3,d:4,e:5}) console.log(value)
× TypeError: ({a:1, b:2, c:3, d:4, e:5})['@@iterator'] is not a functionBut if Coffescript can walk away with its choice then I don't think you'll have anything to worry about.
These are legal CoffeeScript examples:
for key, value of object
if key of object
for value in array
if value in array
for index, value of array
if index of array
The one that's missing is: for value in object
if value in object every element as key, value from arr {
console.log(key, value);
}
every property as key, value from obj {
console.log(key, value);
} every [key, value] from arr {
console.log(key, value);
}
every {key, value} from obj {
console.log(key, value);
}
every [key] from arr {
console.log(key);
}
every {_, value} from obj {
console.log(value);
}
Looks a bit cramped to me but I get the impression that array comprehension is one of the major reasons that people use Coffeescript, so perhaps it's not too much trouble?There's also the issue with some fonts making it hard to differentiate between { and [ but it's an idea that might be worth to think about at least.
fast-each [value] from arr {
}
fast-each {value} from obj {
}
fast-each [value, key] from arr {
}
fast-each {value, key} from obj {
}
The disadvantage is that this looks like array comprehension but isn't.Alternative:
fast-properties key, value from obj {
}
fast-elements index, value from arr {
}
fast-properties value from obj {
}
fast-elements value from arr {
} for (var j=0, len=input.length; j < len; ++j) {
This prevents rechecking the array length on every iteration.You can squeeze out another factor of 2 with typed arrays:
var input_typed = new Uint32Array(input)
exports['numeric typed array'] = function() {
var acc = 0;
for (var j=0; j<input_typed.length; ++j) {
acc += input_typed[j];
}
}
✓ Array::forEach() x 1,999,244 ops/sec ±2.93% (84 runs sampled)
✓ fast.forEach() x 5,161,137 ops/sec ±2.05% (85 runs sampled)
✓ explicit iteration x 27,851,200 ops/sec ±1.32% (86 runs sampled)
✓ explicit iteration w/precomputed array limit x 28,567,527 ops/sec ±1.33% (86 runs sampled)
✓ numeric typed array x 42,951,837 ops/sec ±0.97% (88 runs sampled)
Winner is: numeric typed array (2048.40% faster)
And probably more still if you can figure out the asm.js incantation that makes everything statically typed.The reason it incurs an overhead is because the DOM is traversed every single time the property is read, to ensure nothing has changed. You can see why caching the length is a reasonable optimisation, as you're unlikely to be modifying the collection while looping.
I suspect a lot of the optimization opportunity in browsers today is less about JS or DOM in isolation but more about ways they could work together to improve situations like this one.
(Note: I understand this one is solvable by a proficient/attentive developer, but not all developers are like that, and not all such DOM/JS transition problems are as easily solvable from your JS code)
http://stackoverflow.com/questions/17989270/for-loop-perform...
V8 has excellent profiling tools (exposed in chrome and in nodejs) which should be used first before considering fallbacks. Before seeking a third party library, be sure to check if the function is called many times or is taking a long time.
For example, I found that throwing map out and using a straight array (avoiding the function calls entirely) can be up to 50% faster than using the functional suspects. But that, in the eyes of some people, unnecessarily adds complexity to the code and may not be worth changing
Ironically, that involves certain habit changes that completely obviate the library:
1) avoid map. Just create an array and write a for loop directly at the callsite, putting the code in a block rather than as a separate function
2) avoid indexOf. For single-character indexOf, it's much faster and much more efficient to match the character code (loop and check charCodeAt) than to use the function form
3) avoid lastIndexOf. Same as indexOf, except you loop in the opposite direction.
4) avoid forEach. Learn to love the for loop.
5) avoid reduce. See forEach
Anyone embracing fast.js is sacrificing some performance to begin with.
And with regard to performance, for most Node.js code / webapps if you have a loop with so many iterations that _.map is significantly slower than a for loop then you might be doing something wrong.
It simply refers to 'unrolling' that function call, so a new function doesn't need to be added to the stack.
JS's iteration methods aren't currently recursive, at least in V8[0]. But that's probably due to lack of TCO!
The point of this library is that it provides the same behavior (for 99.99% of cases) and the same abstraction (so same readability / maintainability) as the functions it replaces, but with significantly better performance (for wildly varying degrees of "significantly" depending on usage).
I don't think the idea is that you profile your code and then micro-optimize it using this library. If you're doing that (which you should, for varying definition of "should"), then yes you will want to consider sacrificing the abstraction / brevity for performance.
However, this library seems like a handy way to maintain the abstraction, while gaining performance essentially for free without sacrificing anything worth mentioning. Then you profile (if you were going to / had time to / cared enough to) and optimize from a better starting point. Don't see anything wrong with that.
This is not the case with Fast.js as it breaks away from the language specifications. I'd much rather build, profile, and then optimize the few areas that actually require better performance than potentially introduce bugs by replacing the built-in methods.
HN should have a bot that posts the full quote whenever this line is cited since it's so often abused.
"There is no doubt that the grail of efficiency leads to abuse. Programmers waste enormous amounts of time thinking about, or worrying about, the speed of noncritical parts of their programs, and these attempts at efficiency actually have a strong negative impact when debugging and maintenance are considered. We should forget about small efficiencies, say about 97% of the time: premature optimization is the root of all evil.
Yet we should not pass up our opportunities in that critical 3%. A good programmer will not be lulled into complacency by such reasoning, he will be wise to look carefully at the critical code; but only after that code has been identified."
Plus who said it was / should be used prematurely anyway?
Isn't that essentially what I described? "Before seeking a third party library, be sure to check if the function is called many times or is taking a long time."
On the issue of this particular library: The problem is that you can do much much better by avoiding map, forEach and friends. The overhead of the function calls can be completely avoided by using a direct for loop at the callsite.
So yes, you may find it slightly faster to use this library rather than the native map, but with a small redesign you can avoid map entirely and get a massive performance win.
So I always putting this into not-done-yet category for JS VM related work and it is a very interesting problem to tackle.
I'm sure that is a statistically valid way to measure performance.
So looking at this in Firefox, the "fast" version underperforms the native one for forEach, bind, map, reduce, and concat. It outperforms native for indexOf/lastIndexOf.
In Chrome the "fast" version does better than the native one for indexOf, lastIndexOf, forEach, map, and reduce. It does worse on bind and concat.
It's interesting to compare the actual numbers across the browsers too. In Firefox, the numbers I see for forEach are:
"fast": 47,688
native: 123,187
and the Chrome ones are: "fast": 71,070
native: 20,112
Similarly, for map in Firefox I see: "fast": 17,532
native: 62,268
and in Chrome: "fast": 38,286
native: 19,521
Interestingly, these two methods are actually implemented in self-hosted JS in both SpiderMonkey and V8, last I checked, so this is really comparing the performance of three different JS implementations.bind - fast 2,565,227 ±2.19%
bind - lodash 4,923,373 ±1.51%
bind - native 9,022,396 ±2.12%
on iPad this difference is even larger.
The speed-up for "concat" is surprising to me. I wonder if it holds for "splice" and if that is true across browsers.
exports.forEach = function fastForEach (subject, fn, thisContext) {
var length = subject.length,
i;
for (i = 0; i < length; i++) {
fn.call(thisContext, subject[i], i, subject);
}
};
I knew that forEach was slower than a normal for loop but I was expecting something more. blah | 0; // fast!
Math.floor(blah); // slow(er)! (except on FF nightly)
Caveat: Only works with numbers greater than -2147483649 and less than 2147483648. > -0.5
-0.5
> (-0.5)|0
0
> Math.floor(-0.5)
-1Microbenchmarks like that require a lot of care to measure something correctly. Check out my talk from LXJS2013 for more details: slides http://mrale.ph/talks/lxjs2013/ and video https://www.youtube.com/watch?v=65-RbBwZQdU
But perf problems are more numerous still for these functional methods, because compilers in general have trouble inlining closures, especially for very polymorphic callsites like calls to the callback passed in via .map or .forEach. For an account of what's going on in SpiderMonkey, I wrote an explanation about a year ago [2]. Unfortunately, the problems still persist today.
[1] https://news.ycombinator.com/item?id=7938101 [2] http://rfrn.org/~shu/2013/03/20/two-reasons-functional-style...
It might be that I don't particularly like the language. but it's kind of frightening that we're building the world on that stuff.
* as long as you make it do less work
Or even discuss the fact that if this library is good enough for people to use then maybe the edge cases that the native versions are covering might not be so useful after all?
There is a great, amusing, borderline sci-fi talk by Gary Bernhardt about the future of Javascript and traditional languages compiled to Javascript. My recommendations: https://www.destroyallsoftware.com/talks/the-birth-and-death...
This would likely get even faster if you manually specialized a map/forEach/etc for each caller, but then you might as well just write the loop out by hand: http://rfrn.org/~shu/2013/03/20/two-reasons-functional-style...
On one hand, I'm a big proponent of "know your tools". I'll gladly use a fast sort function that falls apart when sorting arrays near UINT32_MAX size if I'm aware of that caveat ahead of time and I use it only in domains where the size of the array is logically limited to something much less than that, for example.
But on the other hand, I write operating system code in C. I need to know that the library functions I call are going to protect me against edge cases so I don't inadvertently introduce security holes or attack vectors.
If I know that some JS I'm interacting with is using fast.js, maybe there will be some input I can craft in order to force the system into a 1% edge case.
The lesson here is probably "don't use this for your shopping cart", but we need to be careful deriding Javascript's builtins for being slow when really they're just being safe.
It may even be _easier_ to determine the behaviour of these functions than builtins, since determining how fast works just means reading its source whereas to determine builtins' behaviour you must read the ECMAScript specification and then hope[1] that the browser actually matches it perfectly!
[1]: or read the JS engine source, but that's a whole lot more work. :P
Now, if fast.js made sure we had the correct requirements, no distractions, clear milestones, and a dedicated team with no turn over, I'd import it in a second.
I just take objection to the perspective given by the parent post on the _safety_ situation surrounding this library.
[1]: ... not recently, anyway.
I wrote quite a fast "map" (along with the others) that looked a bit like:
exports.map = function fastMap (subject, fn, thisContext) {
var i = subject.length,
result = new Array(i);
if (thisContext) {
while (i--) {
result[i] = fn.call(thisContext, subject[i], i, subject);
}
} else {
while (i--) {
result[i] = fn(subject[i], i, subject);
}
}
return result;
};
I'm not sure if I just used "result = []", but on modern browsers I think that'd be recommended. But yeah, if you're programming for a web browser then using another impl of map is probably going to be a waste of time. native functions often have to cover complicated edge cases from the ECMAScript specification [...] By optimising for the 99% use case, fast.js methods can be up to 5x faster than their native equivalents.The two functions do not necessarily work the same way - the native implementation handles edge cases that the fast version drops.
Generally yes, the results are not expected to be 100% equivalent as these routines do not produce the same results for some inputs.
But as the range of inputs for which the results are expected to be the same is defined, the comparison is meaningful within those bounds (which covers quite a lot of the cases those functions will be thrown at).
It's not like they do apples and oranges.
They do slightly different varieties of apples -- ignore some BS edge cases that few ever use in the real world.
If I rewrite project X into X v2 and throw away 2-3 seldom used features in the process, it doesn't mean that comparing v1 and v2 is meaningless. You compare for what people actually use it for -- the main use cases. Not for everything nobody cares about.
From a programmer perspective, that might not change the fact that it's useful. It has 'meaning' to the programmer in that it helps us solve a particular problem. Typically, app programmers are not as concerned with how the problem was solved.
I like the library, but to avoid this kind of criticism, don't call it a reimplementation. Call it what it is, an alternative library of functions.
I've studied computer science and I don't recall seing any such "formal" definition of meaninglessness.
Nothing in computer science tells you you have to have the exact same behavior in 2 APIs in order to measure some common subset of functionality. You can always restrict your measuremnts to that. Computer science is not that concerned with APIs anyway.
So you can use very formal computer science to compare the performance in big-O of X vs Y implementation of the common supported cases (namely, contiguous arrays).
>This is always possible, but it ignores the original constraints of the problem and hence doesn't shed any additional meaning on how to solve the original problem.
There's no "original" problem that one has to take wholesale. There are use cases, and some care about X, others about Y.
Scientific papers in CS compare implementations of things like filesystems, memory managers, garbage collectors etc ALL THE TIME, despite them not having the same API and one having some more or less features. They just measure what they want to measure (e.g performance of X characteristic) and ignore the others.
This faster version fails on some arrays (any array that isn't contiguous). This is a bug, not a feature. The library user is now responsible for checking that the array is contiguous before using the faster function.
Is it documented somewhere whether array returning/manipulating functions cause non-contiguity?
for (i = 0, total = arr.length; i < total; i++) {
item = arr[i];
}
This will also iterate over gaps in the array. Is this a bug?I'd argue it's actually a future.
For one, it's not forEach for one. It doesn't mask the native implementation. It exists in its own namespace.
Second, it's not like forEach doesn't have its own problems. Like not being backwards compatible to older browsers without shims.
Third, nobody (or very very small values of "somebody") use non contiguous arrays with forEach.
The speed is a great feature, especially if you don't sacrifice anything to achieve it.
>The library user is now responsible for checking that the array is contiguous before using the faster function.
The library user already knows if he is using contiguous arrays or not. After all, it can fuck him over in tons of other cases, even a simple for loop, or a length check, if he treats one like the other. So it's not like he doesn't already have to be vigilant about it.
They say some in their readme. Because the native versions have some extra baggage (sparse array checks, etc).
But to answer in general:
1) JS functions also get compiled to native code for often used code paths. That's what the JIT does after all.
2) A JS function might use a better algorithm compared to a native function.
3) Native is not some magic spell that is always speedier than non-native. Native code can also be slow if it carries too much baggage.
https://code.google.com/p/v8/source/browse/trunk/src/array.j...
These fall outside of the common uses of these methods, and most people don't even know they exist.
http://www.youtube.com/watch?v=65-RbBwZQdU Vyacheslav Egorov - LXJS 2013 talk
i don't know if he actually said these word, but it was the overall theme of this (very entertaining and very enlightening) talk.
> native functions often have to cover complicated edge cases from the ECMAScript specification, which put them at a performance disadvantage.
Aren't these opposing statements?
function add (a, b) {
return a + b;
}
and this in C: int add(int a,int b)
{
return a + b;
}
you're going to get essentially the same performance.The latter sentence you quoted refers to the specific builtin functions that fast.js re-implements. So we first get close to native performance in JS, then we beat the builtin functions by doing less work overall.
A native implementation could have a single flag associated with the array recording whether it is sparse, and use the more efficient code path given here in the common space where it's non-sparse.
Hey I'm not bashing here, I thing it's kind cool for learning purposes attempts to do such thing, but I truly wonder if there is an actual production need for such thing.
lodash is as fast as fast: see the JSPerf results posted elsewhere. John manages lodash well. He has managed it well for quite a while. Why switch to another library with no track record for little or no gain?
What. Is. This. I don’t even.
Why can't we build software on simpler platforms with better implementations?
For example, Haskell provides only a handful of constructs but allows the developer to build sophisticated systems.
Though I agree complex platforms are not necessarily complicated.
I can make the implementation simpler if the underlying specification is simpler. If a language is full of edge cases, it was not designed carefully. If there exists a significantly faster implementation in a significantly slower language that covers 99% of the specification, I'd look for opportunities to simplify.