What causes Ruby memory bloat?
joyfulbikeshedding.com
joyfulbikeshedding.com
Memory managers tend to greedily request a lot of virtual memory because there's usually little harm in doing so, and the act of asking the kernel for more permission (read: more virtual memory) is slow.
Minor quibble: it's not accurate to call `malloc()` and family in glibc the "operating system's memory allocator." That is the memory allocator for C, and Ruby just so happens to be implemented in C. The glibc allocator will use the system calls `brk()`, `sbrk()` and/or `mmap()` to request memory from the kernel (http://man7.org/linux/man-pages/man2/brk.2.html; http://man7.org/linux/man-pages/man2/mmap.2.html). Nothing really changes per the punchline.
Even fluentd had this issue with a default heap size lower than what was needed to initialize itself, causing hiccups, until the GC.stat RUBY_GC_HEAP_GROWTH_FACTOR was tuned in the source code.
Ruby memory bloat is everywhere. Being familiar with gc.stat and being able to tune ruby applications as you test them is a good habit to have if you develop or work with ruby based tools.
Julia Evans, an SRE at Stripe took a sabbatical to work on a ruby memory profiler. Her updates are here: https://jvns.ca/categories/ruby-profiler/
This is a great blog post, it's what I would have wanted to know if I went a bit deeper into this. Good stuff.
To some extent, I agree, at least on 64 bit systems (on 32 bit systems, there was the risk of running out of address space).
However, the memory in question here is, for the most part, not freshly allocated, but was in use once. This means that it used to be backed by physical memory, and before that backing can be withdrawn, the page has to be written to disk.
It seems to me that the best solution would be to call madvise(MADV_FREE) on these regions, in which case they can be unbacked without further ado. I'm somewhat surprised that the memory allocator does not do this itself already.
Because the author does not differentiate between virtual and physical memory, I can’t agree with that. It’s quite possible most of the memory “freed” was never in use. In which case, there’s not much benefit. And there is the probable downside that allocation heavy applications will pay a lot more.
Given the allocation patterns shown, it would seem to me that if two blocks are still in use, it's fairly likely that the region in between also were in use once (there may be exceptions due to pools etc, but generally memory is parcelled out in a linear fashion).
> And there is the probable downside that allocation heavy applications will pay a lot more.
What would the cost be? All that would happen is that the free pages are marked as clean. I'm sure that's not entirely free, but bound to be considerably cheaper than paging the page out and in again.
That's a minor page fault. You get an OS-level exception, switch to kernel mode, process the page fault by marking the page as loaded, then switching back to user mode. That is expensive if you do it a lot.
> but bound to be considerably cheaper than paging the page out and in again.
Yes, of course, a minor page fault is cheaper than a major page fault. But both are more expensive than no page fault, which is what happens if you just never free the page to the OS and there's plenty of available physical memory.
I watched an interview with Bryan Cantrill some time ago and one of the things he mentioned is that libc is considered part of the operating system for other unixes, and Linux is the only one that decided to redefine operating system to exclude libc.
Whatever else ends up there, like POSIX support, is implementation specific.
Regardless, afaik OpenBSD and FreeBSD both ship libc as part of the operating system. It makes a lot of sense, really, because libc is bound to need to make syscalls, which differ strongly depending on the kernel, whereas the user interface of libc is pretty much the same across any operating system.
Yes, that remains very unclear. Only physical memory is interesting in the end, address space exhaustion is not on issue on most systems today and large pages are typically not used. The graphs have "virtual", "dirty" and "clean" in them.
"clean" normally means that the page in RAM is just a copy of a mass storage. If the kernel runs short of memory, it can use this page for other purposes. No harm done, except that the system gets slower when it needs to page in the same page later.
"dirty" normally means that the page in RAM is not backed by mass storage. The kernel must keep it reserved for the current purpose in order not to lose data.
If we assume that author's application does not do swapping (it really shouldn't for 230 MB, otherwise the machine is seriously overloaded and really slow) all heap in use is always dirty. The amount of dirty equals usage of physical RAM.
I am not sure which Linux tool that can easily report the use of physical RAM for the heap. Would RssAnon from /proc/self/status be a good approximation? It certainly contains the stack, too. But that should not grow a lot unless you have infinite recursion.
> Minor quibble: it's not accurate to call `malloc()` and family in glibc the "operating system's memory allocator."
I know. I just explained it like that in order to keep the material digestible to a wider audience.
> I measured RSS with swap disabled, so physical memory usage.
That sounds correct to me. More details you could get from /proc/<pid>/smaps.
There are a couple of other issues I found confusing in your text. If you intentionally make simplifications, I would recommend to add at least a footnote to indicate so.
Linux has only 1 heap. But while the heap is declared to be of a certain maximum size using brk() it does not mean that all of the pages are really in RAM (as before, swapping aside). A page gets only really allocated when it is accessed the first time. And it can get deallocated again using madvise() if the program know that it is no longer needed. malloc_trim() calls madvise(). So you are correct, it does not only move the top of the heap. Probably it was that way many years ago. So in general the Linux heap can be sparsely mapped to RAM. But if you use RSS your measuring takes that into account already.
The arenas are a feature of glibc. They are documented (to some degree) in the man pages, so I would not call them magic. Also malloc_info() prints how they are used. It appears to me that arena 0 is on the Linux heap. If a program is multithreaded, glibc will call mmap() to create an additional memory area to be used for the additional arenas.
While avoiding mutex contention might be a good thing if the threads are allocating many small blocks, the overhead for arenas might not be justified for rarely allocating bigger areas as the Ruby heap implementation probably does it. So limiting the number of arenas might indeed be beneficial, it's just a typical trade-off one needs to understand. Not that that is easy, but magic is the wrong description IMHO.
Let's call the glibc functionality above the system allocator (I'm not sure whether this is the correct name or whether it even has an official name). However, glibc has yet another completely different allocator, the mmap allocator. For allocations bigger than MMAP_THRESHOLD glibc will not use the heap or any of the additional arenas, but just directly make a completely new memory mapping from the Linux kernel. Again these will lazily/sparsely allocated to RAM. The mmap allocator does not use arenas at all.
I am not a regular Ruby user, so I don't know whether the Ruby heap uses glibcs's system allocator or the mmap allocator or what mix of both. (Mixing them is transparent to a program.) However, from your description I would guess that it mostly uses the system allocator, otherwise arenas would not make any change and AFAIK malloc_trim() has no effect on the mmap allocator at all. Have you considered either changing the setting of MMAP_THRESHOLD or the Ruby interpreter code such, that the Ruby heap uses only the mmap allocator?
It appears to me that having the Ruby heap management on top of the glibc system allocator is just not a good idea. The system allocator tries to be a good compromise for a widely varying spectrum of applications. But the Ruby heap management is one very specific case, partially duplicating the work that the glibc allocator does. I'd guess having the Ruby heap running closer to the operating system should improve things. The mmap allocator of glibc might be an easy way to achieve that. Otherwise the Ruby head should use mmap() directly, because it should know best how it wants to use the memory. But that would be much more difficult to implement of course.
P.S. Your visualizer looks impressive. Unfortunately I did not have time to study it in detail. Did you make sure that the code does never access a page that has been mapped but never accessed by Ruby? Because that would dirty the page and increase the RSS.
Dirtying only happens upon write. My visualizer never writes, only reads.
LANG / TIME* / RAM
JS (Node 11.11) / 8.4s / 100Mb
Ruby 2.6 / 19.1s / 14.8Mb
PHP 7.3 / 4.4s / 5.6Mb
Python 3.7 / 24.7s / 4.2Mb
Perl 5.26 / 14.0s / 1.0Mb
*These figures are for runtime, ie. with startup time deducted. Ruby's startup time (0.55s) is much longer than the other languages (Python:0.06s, Perl:0.02s, PHP:0.14s).Ruby's memory usage is 3.5 times that of Python for only a 30% speed gain. Perl uses 1/15 of the RAM used by Ruby and is 35% faster but it could be argued that Perl 5's lack of built-in OO accounts for some of this ..... until you look at PHP which has built-in OOP and uses 2/5 of the RAM used by Ruby whilst performing 3.2 times as fast.
I love that Ruby is designed from programmer happiness but the shine starts to wear off when you look at its memory usage. Slow is bearable as it's only marginal but Ruby's memory usage is orders of magnitude higher it seems. Matz's goal of making Ruby 3 times faster is only half the battle, maybe even only a third. If an increase in speed comes at the expense of even greater memory use then Ruby will not survive.
IO.foreach('logs1.txt') {|x| puts x if /\b\w{15}\b/.match? x } regex = Regexp.new '\b\w{15}\b'
IO.foreach('logs1.txt') {|x| puts x if regex.match? x }
... but performance and RAM usage were identical. pry(main)> 3.times { puts /i_am_a_reg_ex/.object_id }
70131978677720
70131978677720
70131978677720Ruby is just as shiny as it was in 2006, if not more so. If you feel happy writing log parsers in JavaScript or PHP well then all power to you, but I'd rather chew on a shoe.
I do my data processing in Ruby, if there's a performance issue I'll fork out to a little Go widget. But that's when I need a factor ten improvement or more.
var fs = require('fs'), byline = require('byline');
var stream = byline(fs.createReadStream('logs1.txt', {encoding: 'utf8'}));
stream.on('data', function(line) { if (line.match(/\b\w{15}\b/)) console.log(line); });That's interesting to me. I consider bless() and package namespaces to be "built in OO."
Is there some OO concept that combo doesn't provide?
Also, for "Parsing a 115Mb log file for lines containing a 15-character word" ? Wouldn't that just be:
perl -ne 'print if /\b\w{15}\b/'
I'm surprised that's not faster.
If blessing a hashref is real OO why has so much energy been expended on Moose and its offspring? Before I left, back in 2013, there was also much gnashing of teeth over whether Perl needed a MOP so I don't think everyone agreed that bless() is enough.
That's a good question. It's always mystified me. A little boilerplate doesn't bother me. As far as I can tell, it's because people didn't like writing the constructor boilerplate and getters/setters. Maybe for the "isa" pretend types?
https://stackoverflow.com/questions/15529643/what-does-mallo...
That isn't free (it requires cache & TLB space), but it's entirely possible that Ruby is slow enough on the interpreter and data model level for this to not matter much.
Edit: Turns out malloc_trim doesn't actually modify the mapping but rather uses madvise(DONTNEED), so a higher address resolution cost probably only materializes under memory pressure.
All the experts say "oh, Ruby uses lots of memory for [reason] and it can't really be fixed", so no one even tries.
Until someone comes along who is either motivated, smart, or ignorant(!) enough to try to fix it anyway, and finds that the commonly accepted answer was wrong.
This happens all the time, especially in science. Trust, but verify, I suppose.
This isn’t true at all. It’s well understood that jemalloc 3.x exhibits lower resident set size because it more readily releases pages.
This idea has been around for at least 3 years: https://bugs.ruby-lang.org/issues/12236
My current employer typically uses M5 instances which have a ratio of 1 vCPU : 4 GiB Ram.
Running Unicorn you'll probably only want 1.5 instances per vCPU. Even a memory heavy Rails app is probably only going to utilise ~20% of the available memory.
Running threaded Puma, you probably want only a single process per vCPU and maybe 5-6 threads. In my apps running 5 threads per process typically increases memory of the process by 20%. So in that instance you'd only utilise 15% of the available memory on a M5 instance.
If you are having memory issues on Rails, then quick wins are upgrading your Ruby version. I saw 5-10% drop in memory usage with each of the major version 2.3.x -> 2.4.x -> 2.5.x.
Also if it is an old app, check you've not built up cruft in your Gemfile. Removing unused gems can be another quick win for reducing memory usage.
* prove it is really the issue
* fix it
* prove that you fixed ita simple multithreaded HTTP proxy server written in Ruby (which serves our DEB and RPM packages)
then I would reach for Linux ipvs or haproxy, and apache or nginx to do the serving. Good tools already exist for these things, it's a shame not to use them.
(And we have a Ruby dev group, so please don't accuse us of having a phobia or hatred of Ruby.)
While I agree with you, I also see no harm in experimentations and changes to the status quo for the sake of trying to build something better.
In any case it's just something thrown together to solve a need quickly and effectively. The performance characteristics might not even have been an issue, they just caught his eye.
Re-inventing the wheel is almost certainly harder than learning the configuration syntax of any of these.
Writing a memory visualiser is exactly what I needed and planned to do next, so this is a great contribution for anybody working on problems like this.
Java has a compacting garbage collector. So it should not be affected.
Best thing is this can be called directly from Ruby with FFI.