How PostCSS became 1.5x faster by changing 2 lines of code
evilmartians.com
evilmartians.com
Do not even try to write what you think is effective code before
benchmarking. The VM has many clever optimizations. And even if
you will somehow learn of all of them, in the next release they
still can be changed. Instead, write a simple and clean piece of
code, make a benchmark, find the real bottleneck and rewrite code
in small parts.
Every now and then I see some strange code at a client I work with, and the justification is "Performance". Because somebody thought it might be faster - But they didn't check it back then. And they don't have an automated performance test that continuosly validates if their "dirty but fast" code is still faster than the clean version...Don't even ask how many hours of developer time I saw wasted because of "dirty but fast" code. Where the code often was not much faster then the clean version. Or it was faster, but speed was not crucial in the area where the code operated.
Also, I guess a lot of the optimizations from old blog posts or the first "Effective Java" will not yield amazing results anymore, as compilers and runtimes are getting better.
I also like this one:
Do not think that programming in C++ or any other lower-level language
is a must for having good performance. Good architecture, benchmarking
and profiling are far more important.
But I don't think it is as clear-cut as the author writes it. For some problems, having total control over your memory layout and what gets exectued when (i.e. C or C++ or ...) can lead to huge benefits. And sometimes, the runtime is just more clever optimizing the code than you are.As any follower of @mraleph (http://mrale.ph / https://twitter.com/mraleph) is aware, it's very easy to benchmark the wrong thing (or benchmark more or less nothing), or to write an isolated benchmark which doesn't reflect real-world use.
With a bit of thought this even leads to cleaner code: do you really need to pass a large chunk of data around when you only need a particular precalculated quantity?
But generally benchmarking and applying Amdahl's law is the way to go. First identify which part of the system is slow!
I think there is some opportunity here to improve compiler output. I know that GCC can already explain why it doesn't vectorize loops (-ftree-vectorizer-verbose=2), but many other optimizations without such output can still make or break the performance of your program.
delete this.indexes;
wouldn't the author avoid the performance regression with this.indexes = 0;
and still avoid external state?Although, you might say that the previous version is 1.5 times slower than the current version.
"33% times faster" doesn't make sense. You need to drop the "times" or use "0.33 times" for that to make sense.
> But if V8 thinks that the code is “too tricky”, it keeps it in slow but dynamic form.
The next paragraph shows what caused V8 to not optimize :
> For example, in the first set of changes I’ve used a global variable defined outside of the class, so it had an external state. In the second one, I’ve changed the class structure on the fly. As a result, V8 did not compile it to effective typed native code.
But you're right, saying "it's probably that" is not the same thing as delving in the code.
If you want gory details I recommend Vyacheslav Egorov's blog (mrale.ph). Here's an old slide of his that lists all the ways you can force an object into dictionary mode (at least ca. 2011):
Using Object.seal, Object.freeze and Object.defineProperty with writable, enumerable, configurable not set to true does not cause object named properties convertion to dictionary mode. However they still convert elements storage (one containing properties with names "0", "1", etc) to dictionary mode.
Similary accessors don't cause object properties storage conversion to dictionary mode unless there is a transition clash.
var obj = {};
obj.__defineGetter__("foo", function () { return 0; })
print(%HasFastProperties(obj)); // => true
var obj = {};
obj.__defineGetter__("foo", function () { return 0; }) // Transition clash
print(%HasFastProperties(obj)); // => false
Nothing changed with respect to `delete obj.foo`: this always converts named properties storage conversion to dictionary mode. However `delete arr[index]` doesn't (at least for arrays that were not in slow mode already), rules for objects depend on the amount of holes in the elements part of the object.The usual disclaimers of "profile it first" and "YMMV" apply, but it's useful to be aware of potential low hanging fruit when JS performance is critical to your application.
But I'm guessing it didn't help for this particular problem. Deleting properties from an object can hurt performance in v8, but it's not the deletion that's slow, it just makes the later accesses to that object slower. So the line of code that needed changing may not have been anywhere near the code that profiled poorly.
> never optimize, always profile
Optimization, when done correctly, is a feedback loop between your assumptions and the profiler. Neither one of these, on their own, is enough -- your assumptions can be wrong, but the profiler has an extremely poor signal to noise ratio. (Note: Profiling wouldn't have caught this issue, since the problem spot was not actually slow, it just made everything else slower)
I work in game development (not primarily HTML5, although I have done HTML5/WebGL projects), and performance is a hard requirement for me. Not making 60fps reliably for the hardware we need to support is a show-stopping bug. The only way to avoid this reliably without needing to rewrite sections is to design with the optimizations you might need to perform in mind. Putting optimization off until the end would be irresponsible.
> PostCSS, libsass and Less are already fast enough for any real-world task. Running benchmarks like these is like listening to audiophiles comparing gold-plated cables for their hobby systems. Don’t use these benchmarks as the main criteria for decision making when you are choosing a tool.