When allocators are hoarding your precious memory (2021)
algolia.com
algolia.com
It turned out they also had lots more CPU cores (they were repurposed former DB machines), and by default the memory manager (of the JVM, I suppose?) scaled the allocation block size the number of CPUs, so it tried to allocate far more memory than necessary, and since this was still a 32-bit JVM it quickly ran out of space for its allocations.
Once we found out what the issue was, it was pretty easy to tune with an environment variable.
I know in the grand scheme of things it’s meaningless, this is just one of those things that’s like someone scratching their fingers down a blackboard when I see it day to day.
Tangentially, "Prod" is justified since it's being treated as a proper name, like "Paris", "Roderick", or "Mount Rushmore."
But yes, "PROD" is harder to explain, and I suspect it might be bleed-over from looking at code or config files where it gets capitalized.
Think about book titles. "Exploration of Paris", the important words are capitalized.
I can imagine many people capitalize the environment names without thinking too hard about _why_ they do it.
This is, of course, descriptive typography. Prescriptive typography will vary from style manual to style manual.
Well that's the dictionary definition anyway. In practical terms, I think most people use acronym as a synonym of initialism, at least in the US. This is one of the places where the dictionary definition doesn't match with real world usage.
I'm trying to train myself out of this nonsense, but some habits die hard.
Not sure if this is a bash hold over to differentiate ENV_VAR vs local reg_var or like an os/global constant naming hold over.
Generally if you call "env" on your cmd line (all of my)/most keys are ENV_VAR capitalized.
We ended up moving to tcmalloc, mostly because the Debian packages worked the best for us. we did testing between the different malloc replacements, and honestly, there wasn't a huge difference between them for our workloads, although they _did_ have differences, with minor variations in peak/avg memory usage and CPU usage.
Where we are now is better, but has its own problems.
Not really. Lots of rust users come from high level languages so when people come in and complain about rust being unexpectedly slow allocation issues are in the top 5 at least. It’s incredible how shit the standard allocators are in almost all systems.
This is not inevitable, platforms could provide allocators which are less awful. Obviously they can’t be as fast as specialised runtime facilities but when you see the gains many applications get by just swapping in jemalloc or some such…
It’s not the beginner’s fault that most system allocators suck so hard you must do your utmost to limit their use.
Many of the features in modern programs didn’t exist in the 90’s not because they hadn’t been thought of but because people spent all their energy getting the first fifty features to work reliably. Every app had a couple cool features. Most of them are de rigeur now because we can.
We're still waiting on that to happen.
There's a difference between a developer who understands manual memory management choosing to use a tech that does it for you and a developer choosing to use a tech that does it for you so they don't have to learn manual memory management.
manual memory management is table stakes for any halfway competent developer.
That said, there are a few reasons beyond just that. For the primary sources: https://github.com/rust-lang/rust/issues/36963
Huh. I’ve seen lots of people write slow Rust code because they didn’t realize there were allocations. But the complaint is usually “my rust program is only 5x faster than my equivalent python program instead of 100x how come”?
Can you point out some good blog posts about Rust or C++ being slower than a “basic runtime” due to inferior allocator?
Escape analysis as you’re alluding to isn’t needed in this model because the amount of times this helps you (ie you put it on the heap but the compiler can figure out it can live on the stack) is about 0. You need escape analysis in managed memory languages where everything is nominally a heap allocation and the compiler is responsible for clawing back performance through escape analysis.
Are too
> They are supposed to be general-purpose allocators with requirements that may not match your workload.
I don’t think they match any workload which exists anywhere anymore. Possibly their workload existed back in the 90s.
Jemalloc is also a general-purpose allocator, it‘s the one freebsd uses.
> FUD around switching to something like mimalloc or the most recent tcmalloc
Certainly, and the specific FUD about libc++ is unwarranted for tcmalloc, considering the developer of tcmalloc also uses libc++.
As for what allocator you should be using, it depends on your workload. You'd need to test out some different options to see what works best. And "best" might also be something you have to define. Maybe you're ok with higher baseline memory usage if allocations are faster. Or maybe you want the lower memory usage.
Most workloads don’t care of course, but glibc’s penchant for hanging onto RAM when not needed is one very user visible consequence of this.
I am honestly surprised no one has bothered replacing the system default for distros as a better default allocator would free up RAM since both tcmalloc and mimalloc do a good job of knowing when to release memory back to the OS (not to mention that they’re generally faster allocators anyway).
First off, GCs are quite deterministic, unless they are calling rand() someplace. I've walked through GC code under controlled conditions with identical repeated allocations, GCs do the exact same thing under the exact same circumstances every time.
Second off, you have to pay the price of allocation and deal with memory fragmentation somewhere. GCs typically pay the price at the tail end (deallocation, compaction) but have absurdly fast allocators (a handful of instructions), where as manual memory allocators pay the price for allocation (finding a slab of memory) up front, and unless you know how to work around the details of your allocator, you end up dealing with fragmentation yourself.
FWIW this also means many, many, types of applications can gain a boost from intelligent memory allocation strategies. There is the famous example of the cruise missile that never worried about deallocation because it would explode before it ran out of memory, but scenarios involving bespoke memory management being better than "just use libc!" are actually quite common.
Have a stateless microservice? If you know the max size of memory it'll use, for each connection that comes in just allocate a slab of memory and when the connection closes free the entire slab, far more efficient than using a GC or any standard allocator. Also having to calculate the max amount of memory a single connection can use is a great way to figure out how much load an instance can handle (depending if you are CPU bound or not).
For certain stages of compilation, compilers can get away with not freeing memory and just exiting and letting the OS reclaim everything. (Of course less doable now days with compilers running as background services recompiling code as it is typed, and also whole program optimization has made it so compilers use a lot more memory, so I don't know if this strategy is still in use in these modern times!)
The tl;dr is that freeing memory is complicated no matter if you have a GC or not. It is something every developer should be thinking about from time to time.
Is that wrong? Or obsolete? Or GC-dependent?
But also, just because something is theoretically deterministic (e.g. it might technically be possible to work out when a GC sweep will run, and multiple runs of the exact same program will cause the exact same set of GC sweeps), actually working out when a GC sweep will run is so non-trivial the only way to actually work it out is to run the program and watch it happen, and also it becomes impossible to predict how any changes at all to a program will affect how the GC runs as a result (other than making the change and running it).
Or are there ways to deduce when GC sweeps will run for a given program, without actually running it?
GC-dependent.
Though most GCs are threaded so there is always a level of indeterminism there. That being said, the same can be said of any application. Unless you are doing actual real time programming you always have at least a little indeterminism for how your app will behave based on the rest of what's going on in the OS.
> you can't control exactly when the GC is run?
Depends. A number of SDKs will expose a "trigger the GC" API, though that's generally seen as an anti-pattern.
But speaking of GCs, there's roughly 2 big categories of GCs. The first is a full stop the world while the GC is running, the second has the GC running concurrently with the application.
For a full stop the world GC, that is fairly predictable when it comes to how it will behave. GC is pretty much always triggered on an allocation failure. For full stop the world GCs, that will involve tracing through the live memory roots and somehow marking what's still used and discarding what isn't (State of the art, AFAIK, is a moving collector where live data is moved to a new memory region. This has an advantage of always compacting the data can (but isn't always) be pretty cache friendly as related data is very often colocated.
For the second type, generally what's happening is some signal is sent that indicates it's time to collect memory. The world stops for that signal to propagate but after that happens the app is allowed to continue to run while the garbage collector runs on a parallel thread. The whole app is only completely stopped to wait on when you run into a "you are still collecting and I want to allocate, but there's no space" so of scenario. In these second types of applications it's common for there to be an added GC logic hoisted into the application when memory is accessed because "obviously this is still live so do the GC work as part of the regular app work". This sort of collector would be the hardest to reason determinism about as what happens when you access a memory location will be determined by where the collector currently is and what the application is doing.
The first type of collector can be seen in the JVMs parallel and serial collectors.
The second type can be seen in the JVM's G1GC (partially), ZGC, Shenendoah. It's also fairly closely related to how Go's GC works and how (AFAIK) javascript's GC works.
There is 1 other semi common category of GC, that's reference counting. There you have a VERY high level of determinism as there's no side threads and collection happens purely based on application actions and not any sort of opaque memory investigation. However, RC has it's own set of issues that tracing collectors do not.
Which is the case for memory fragmentation without a GC.
A naively written program in C++ will, if left to run long enough, eventually fragment memory so much that allocations start to fail.
But before then you'll have "non-deterministic" behavior of the allocator jumping around the heap trying to fulfill allocation requests.
I don't know how GCs work. Don't they still need to find free memory somehow during allocation? I always thought of GCs as manual memory allocators + automatic deallocations
There are some GCs that are much more complicated than that. However, typically, especially for new small allocations, it's just a bump allocator. For larger allocations or when talking about being a generational collector that's when things get more complicated.
For JVM's GCs, most of the collectors are "compacting" which means that after a collection the region collected is left with all the live data unfragmented. That means allocating to a given region is simply a process of keeping track of the pointer for that region and bumping it when something needs to go there.
The JVM did have the now defunct (thank goodness) CMS collectors which would still compact the minor collection regions, but for the old heap regions it'd maintain a skip list. It would only compact the old gen region when after a major collection it still failed an allocation. In that case it'd take the swiss cheese old gen and squish it together.
Go does not have a moving collector so I'm guessing it is doing something like an arena allocation (perhaps just relying on jemalloc and calling free at the right times?). In which case go's allocation speed will be around C's speed. They could be doing something different but I didn't stumble across that googling.
And yes, I understand that technically glibc is not a part of "Linux", but...