JavaScript- Lodash vs Js function vs. for vs. for each
github.com
github.com
Even on one of the first lines, there was already a mistake with how reduce is called:
let avg = 0;
console.time('js reduce');
avg = posts.reduce(p => avg+= (+p.downvotes+ +p.upvotes+ +p.commentCount)/3,0);
avg = avg/posts.length;
console.timeEnd('js reduce')
The way he's using it makes it no different from a forEach. It should be
`posts.reduce((p, accumulation) => accumulation + blah, 0)` instead.Also the timing isn't very sophisticated. Doing microbenchmarking right is basically black magic but the mistakes he's making are really basic (only running each version of the code once, not preventing dead code elimination, etc). Someone mentioned elsewhere that the timing for the large case is faster than the small and my guess is that the entire calculation got JIT'd out since it's unused.
The difference with a simple for() are quite significant.
The data collected is not on running it once, the results are avg. of running same code at least 15 times.
It works on `arguments`, it works on strings (iterating through characters), it works on objects (passing the key as the iteratee's second parameter in place of an index value), it works on HTMLCollection pseudo-arrays. It also doesn't throw on null or undefined.
Lodash may be fast, but recently I've been avoiding the basic "lodash function with native js equivalent" for one particular reason: stepping into js native functions when debugging (node inspect) is a breeze, and a complete nightmare when using lodash.
[...iterable] Array.from('string')
[...'string'] Array.prototype.slice.call(htmlCollection);
Which has been the canonical way to cast them. But it will still throw on null or undefined.The issue was that the collection being iterated over was a non-generic .NET 1.0 DataTable. Using a foreach loop would implicitly box and then re-cast each object, while the for loop directly accessed the correctly typed .Item() and did not need to do that.
Ironically, the body of the loop was a fairly tricky logistics algorithm I had just written, so I had every reason to assume the problem was on the inside. Imagine my surprise when I changed it to a for loop - strictly to access the index and print out some timings - and watched the procedure suddenly become instant...
Well, yes, but the relative difference between a for and foreach loop is miniscule - not 50x difference. In absolute terms, the overhead of both is barely measurable, and is extremely unlikely to be a bottleneck in any client-side application.
For example, here are the results for the for loop 'reduce' on the small data set:
100 items - 0.030
500 items - 0.574
1000 items - 0.074
That doesn't make sense to me. How can a reduce on 1000 items be drastically less than on 500 items. Unless I'm misunderstanding something, I can only conclude that it's either A) a typo or B) they only ran this test once and the random data for the 500 test was exceptionally tough to deal with.
Either way, I would love a little more detail with the data before I trust it.
EDIT: Maybe what happened is that the 500 case made the engine conclude that the function is called frequently enough to be optimized, and the 500 run both spent a while with the slow version and spent some time optimizing it, while the 1000 case got to exclusively use the optimized version.
Just goes to show that benchmarking complex multi-stage JITs can be hard.
There are a ton of things one needs to do to make reliable microbenchmarks (make sure the functions are getting optimized, make sure they're not eliminated as dead code, etc.) and I don't think this repo does any of them.
If using node, use process.hrtime() rather than console.time()/timeEnd(): https://nodejs.org/api/process.html#process_process_hrtime_t...
Then, computing the length multiple times is not a good idea. You should save the length in a variable:
// no
for(let i=0; i<posts.length; i++) {
// yes
for(let i=0, n=posts.length; i<n; i++) {
Finally, it is not recommended to analyze performance in this manner. A slight little change elsewhere in your program can affect the performance very abruptly.This is because the gatekeepers of performance are: inline caching, hidden classes, deoptimizations, garbage collection, pretenuring, etc.
The quick microbenchmark I checked this on: https://jsperf.com/for-to-length/1
// yes
for(let i=0, n=posts.length; i<n; i++) {
There's usually no need to cache the length of the array this way. Modern JS VMs are plenty smart enough to do it automatically (unless there's code in the loop that looks like it might change the array length).Working on it.
That said, it's still a good idea, even if it's just for pointing out that the constraint on the loop won't change.
However, you are right about the performance benchmarking factor. Good news, I have done analysis on the inline cache, warm cache and working on how to get GC in place and hidden classes to get better results.
Although Ramda has forEach, I augment it with a version of each(func, data) where data can be an array or map, and func(val, key) where key is the key of the map item, or the index of the item in the array.
I feel this abuse of notation makes for more readable / smaller / uniform code [ having no explicit for loops ]. Also takes less conceptual space.
https://github.com/lodash/lodash/commit/6b2645b3106b0ed9ebec...
Does the first algorithm get an unfair cold cache disadvantage?
Take a look at the array size 500 results for example. It is slowest for Reduce. Or take a look at 5000. There it is second fastest for Reduce, and slowest for Map, Filter, and Find.
There are cases where large data sets are generated in the browser, but that is not even close to often.
Even in the world of game dev, JavaScript is not common at all.