Fall cleaning: Optimizing V8 memory consumption
v8project.blogspot.com
v8project.blogspot.com
1) Reducing the V8 heap page size from 1M to 512KB ... results in 2x memory reduction
2) The memory visualization tool helped us discover that the background parser would keep an entire zone alive long after the code was already compiled. ... resulted in reduced average and peak memory usage.
3) C++ compiler not packing structs optimally for V8s case - manual packing saves some peak memory.
So C/C++ compilers allow you turn this off either by specifying the data length yourself (bit fields) and/or specifying attribute(__packed__) directive in the case of GCC for example. In this case the packed directive resulted in suboptimal packing so they must have gone with bit fields I suppose to do manual packing.
In my case I'm tuning a 3D game, and the heap usage is a sawtooth pattern with GCs about once per second. Not awful, but it would be nice to smooth it out. But playing with Chrome's various memory profiling tools I've never been able to discover where the allocations are coming from - the results always seem to be dominated by apparently internal stuff (like entries called "(code deopt data") or such). Does anyone know any good techniques for such things?
Actually now that I take a second look, some relatively simple demos of the same engine (Babylon.js) show the same sort of behavior. Some rather trivial three.js demos do as well. I might be dealing with something that's just a fact of life for webGL rendering.
Random example (not mine) showing a vaguely similar sawtooth of heap memory usage:
I'd love to take a look deeper - contact info is in my profile if you're interested.
One suspects that DevTools' memory profiling should be the place to start, but I haven't found any way to get it to shed any light on where allocations are occurring. So that's what I'm asking about here.
(I mean, not to suggest that what you describe isn't useful - I'm just hoping to use profiling tools to attack from a different direction.)
>> not so easy to jump into a three.js-sized chunk of 3rd party code and find atomic things that can be turned off
Incidentally, doing a dissection like this is actually an engaging way to learn a big chunk of 3rd party code compared to just staring at it. By stubbing out various pieces you learn where the joints are where things are entangled nests of sorrow.
With that said, do you have any advice on how to use it in practice? Timeline profiles tell me my app goes through maybe 5MB of heap per second, but when I use this feature for say, 5 seconds, it tends to report 2-3 functions as having allocated 16kb each. (And if I run it again, I get similar results but with a different 2-3 functions.) Is it just reporting a very small subset of allocations?
If you want to look at all the allocations during the interval, then you can use the 'Allocation timeline' profile – this will give you all the allocations but note that this might have significant overhead.
* https://gist.github.com/krisselden/d3ce3cbb37cc6035b0927fdbf...
It flips some handy flags providing useful output, this output can quickly illuminate issues the regular tools do not (yet).
Running this on the example you linked to bellow, shows that a series of functions are deopting and optimizing repeatedly. most likely causing at-least some of the sawtooth pattern you see.
example output (likely related to the problem):
```
removing optimized code for: r.getViewMatrix]
[removing optimized code for: r._isSynchronizedViewMatrix]
[removing optimized code for: r._computeViewMatrix]
[removing optimized code for: r._isSynchronized]
[removing optimized code for: r._computeViewMatrix]
[removing optimized code for: r._isSynchronizedViewMatrix]
[removing optimized code for: r._isSynchronized]
[removing optimized code for: r._computeViewMatrix]
[removing optimized code for: r._isSynchronizedViewMatrix]
[removing optimized code for: r._isSynchronized]
[removing optimized code for: r.getViewMatrix]
[removing optimized code for: r._isSynchronizedViewMatrix]
[removing optimized code for: r._computeViewMatrix]
[removing optimized code for: r._isSynchronized]
[removing optimized code for: r._computeViewMatrix]
[removing optimized code for: r._isSynchronizedViewMatrix]
[removing optimized code for: r._isSynchronized]
[removing optimized code for: r._computeViewMatrix]
[removing optimized code for: r._isSynchronizedViewMatrix]
[removing optimized code for: r._isSynchronized]
[removing optimized code for: r.getViewMatrix]
```
---note: Credit for this should go to @krisselden the author of the above gist not me.
For this reason I always let the game run for 10-15 seconds before I profile, figuring that by that time most of the opt/deopt churn will be finished. And this (very useful) script seems to back this up - I get output like yours at startup, but if I wait a while and then rm the output files, further output is relatively minimal.
So I'm inclined to think that opt/deopt stuff isn't the issue, and there really are lots of JS heap objects getting allocated somewhere. At the same time though, when I used Chrome's built-in memory profiling I see a bunch of deopt-related strings, so maybe I'm way off base. If anyone sees what I'm missing please do clue me in.
(Also: great tip on the script!)
In my experience, these add up real quick and are often indicators of a larger "instability" issue that remains well after the "deopt churn" appears to settle, but continues manifests in the form of some heavy GC.
Note: many internal structures related to the JIT (IC/hidden classes/code gen etc) can themselves cause sufficient GC pressure, as can the code you described as "de-opted permanently".
Interestingly it is also possible for the above mentioned GC pressure to itself cause some fun (even more GC pressure): https://bugs.chromium.org/p/v8/issues/detail?id=5456
This may not be the root of your issue, but I would be careful to rule it out entirely to quickly.
Anyways, best of luck!
"Reducing the V8 heap page size from 1M to 512KB results
in a smaller memory footprint when not many live objects
are present and lower overall memory fragmentation
up to 2x."
Is it common to say something's shrunk by 2x? Why not say 0.5x (or 50%, half, etc.) I understand growth of 2x and assume this is a mistake, though I'm open to convention.I.e. shrinking 2x and then growing 2x would bring you to the original.
Though in this specific case, I would have simply said "halved".
[1] http://kinoma.com/develop/documentation/technotes/introducin...
EDIT: Kinoma claims ES6 compatibility is at 98% as of 2016-01-02.[2]
[1] https://kangax.github.io/compat-table/es6/ [2] http://www.kinoma.com/develop/documentation/js6/