Linux ate my RAM (2009)
linuxatemyram.com
linuxatemyram.com
Linux, by default, is making the very reasonable assumption that the marginal cost of converting empty physical memory into caches and buffers is very near zero. This is fundamentally reasonable, because the cost of converting empty memory into used memory isn't really any cheaper than converting a clean cached page into used memory. It's a little more subtle when you take accounting into account, or when you think about dirty pages (which need to be written back to clear memory), or think about caches, but the core assumption is a very reasonable one.
Except for on some multi-tenant infrastructure. Here, "empty" pages don't really exist. There's mostly not an empty page of memory kicking around waiting (like there is on client devices). Instead, nearly all the memory on the box is allocated, but each individual guest kernel doesn't know the full allocation. In this world, the assumption that the marginal cost of converting empty to full is zero is no longer true. There's some real cost.
Projects like DAMON https://sjp38.github.io/post/damon/ exist to handle this case, and similar cases where keeping empty memory rather than low-value cache is worse for the overall system. These kinds of systems aren't super common, especially on the client side, but aren't unusual in large-scale cloud services.
Gray and Putzolu's classic "The 5 minute rule for trading memory for disc accesses" (https://dl.acm.org/doi/pdf/10.1145/38713.38755) from 1987 is probably one of the most important CS systems papers of all time. In it, they lay out a way of thinking about memory and cache sizing by comparing the cost of holding cache to the cost of access (this isn't the first use of that line of thinking, but is a very influential statement of it). Back then, they found that storing 4kB in RAM for 5 minutes costs about the same as reading it back from storage. So if you're going to access something again within 5 minutes you should keep it around. The constants have change a lot (RAM is way cheaper, IOs are way cheaper, block sizes are typically bigger) since then, but the logic and way of thinking are largely timeless.
The 5 minute rule is a quantitative way of thinking about the size of the working set, an idea that dates back at least to 1968 and Denning's "The working set model for program behavior" (https://dl.acm.org/doi/10.1145/363095.363141).
Back to marginal costs - the marginal cost of converting empty RAM to cache is zero in the minute, but only because the full cost has been borne up front when the machine is purchased. It's not zero, just pre-paid.
Running the numbers - assuming 4k record size instead of 1k, ignoring data size changes, ignoring cache, ignoring electricity and rack costs, selecting a $60 Samsung 980 with 4xPCIe and a $95 set of 2x16GB DDR5-6400 DIMMs...I get $0.003/disk access/second/year and $0.0000113 for 4k of RAM, a ratio of 264.
That is remarkably close to the original paper's ratio of 400, even though their disks only got 15 random reads per second, not 20,000, and cost $15,000, and their memory cost $1000/MB not $0.002/MB.
I'm not sure the "Spend 10 bytes of memory to save 1 instruction per second" works equally well, especially given that processors are now multi-core pipelined complex beasts, but working naively, you could multiply price, frequency, and core count to calculate ~$0.01/MIP (instead of $50k). $0.01 is about the cost of 3 MB of RAM. Dividing both by a million you should spend 3 bytes, not 10 bytes, to save 1 instruction per second.
If this is a Hetzner machine then yes, but enterprise SSDs costs more, especially from enterprise vendors. But this only drives the storage cost up.
More so, if you tend to send some big amount of data every 5 minutes and you are somewhat constrained by memory (32 / 1000 x 100 = 3.2%) then it would be easier to just read it from the storage again. If you are not constrained by storage bandwidth, of course.
And by the way, the latest gaming consoles (at least PlayStation?) is designed around this concept - they trade having big amount of RAM (which in case of PS5 is shared between GPU and the OS) to just loading assets from the storage extremely fast 'just in time'. Which works fine for games.
Or, in other words, you get to fully use what you paid for.
I was under the impression that at least in some virtual machine types the guest kernel is collaborating with the host kernel through vm drivers to avoid this problem.
But I do vividly remember the help desk trying to figure out one last issue. Some process was consuming all my resources.
They never could figure out why "System Idle Process" kept doing that.
I couldn't answer that at time. Took a bit more years and understanding until I remembered that situation and it was obvious for me.
CPU Time is a metric of how much 'real time' was spent on the process, so a single thread running 100% time [on a single core with no HT] would give you a full 'real time' second. System idle 'consumes' all the idle CPU time so if the system is doing nothing it's get a second of CPU Time per execution thread. And there are multiple execution threads on almost anything later than 2005. On my T440 with i3-4010U (2 cores, 4 threads) the System Idle consumes ~3 seconds of CPU Time per second, because there are ton of shit running in the background so the system is never 100% idle.
We'd demand standard alarms for things like memory leaks / out of memory conditions / high than normal memory usage, as to get 99.999% uptime we want to be paged when problems like this would occur. Except a bunch of platform did the extremely naive implementation and included recoverable memory in their alarm conditions. So inevitably someone would log in and grep the logs or copy some files to the system, and hit the alarm conditions.
And there were some vendors who really didn't want to fix it, they would argue that recoverable memory is in use, so it should really be part of that alarm condition.
The memory access pattern was pretty much pessimal for my use of the box as a workstation. I'd use it from 7am -> 8/9pm every day, then when I'd walk away from the keyboard, I'd watch HD recordings (which could be 7GB or more per hour). Those would get cached in memory, and eventually my workstation stuff (emacs, xterms, firefox, thunderbird) would start to get paged out. In the mornings, it was painful to start using each application, as it waited forever to page in from a spinning disk.
I eventually wrote an LD_PRELOAD for the DVR software that overloaded open, and added O_DIRECT (to tell the kernel not to cache the data). This totally solved my problem, and didn't impact my DVR usage at all.
It's ok, no shame. But as I understand it, FreeBSD would prefer to throw out (clean) disk cache pages under memory pressure until somewhere around FreeBSD 11 +/- 1, where there were a few changes that combined to make things like you described likely to happen. Heavy I/O overnight might still have been enough, and I'm not going to test run an old OS version to check ;)
I can't find the changes quickly, but IIRC, older FreeBSD didn't mark anonymous pages as inactive unless there was heavy memory pressure; when there was mild memory pressure, it would go through the page queue(s) and free clean disk pages and skip other page; only taking action on a second pass if the first pass didn't clean enough. This usually meant your program pages would stay in memory, but when you hit memory pressure, there would be a big pause to mark a lot of pages inactive, often too many pages, which would then get faulted back to active...
Current FreeBSD marks pages inactive on a more consistent basis, which is nice because when there is memory pressure, chancws are there's already classified pages. But it can lead to anonymous pages getting swapped out in favor of disk pages as you described; it's all tunable, of course, but it was a kind of weird transition for me. After upgrading the OS, some of my heavy i/o machines saw rising swap usage running the same software as before; took a while to figure that out.
- heavy disk access because we were writing real-time images to an SSD at about 500MB/s
- our application was steady-state about 4GB of RAM and we had 32GB available on the platform
- the serial port that we received data from was, under the hood, using DMA
In certain cases, Linux would completely run out of free pages (28GB of it being used for cache on files we were never going to read again). These were all available pages but just occupied at the exact moment. The serial driver would request a page for DMA when it received an interrupt and being inside an interrupt context would request that page with NOBLOCK. That meant that kmalloc would return NULL instead of giving a page, since it would need to evict one of the cache pages before one was available. The serial driver would then blow up and never retry the DMA transaction.
Fun to debug that one!
(This was of course because of having too many apps relative to my RAM. Not because of disk caching.)
The OOM behavior is not pleasant for a desktop system.
I remember some time back there was discussion about improving the OOM killer, but I don't know what came out of it.
I've heard some good results with it and the applications locked in memory is configurable.
Out of memory: kill process 12345
Killed process 12345 (sshd)
is the funniest and ugliest message to see on the iLO/VM console. cat > /etc/tmpfiles.d/mglru-min-ttl.conf <<EOF
w- /sys/kernel/mm/lru_gen/enabled - - - - y
w- /sys/kernel/mm/lru_gen/min_ttl_ms - - - - 1000
EOF
and reboot.
I've been struggling with this issue, as many others, for years. Now I can run two VMs with 8 GB physical RAM and 5+ GB swapped, and it barely noticeable.More information, although a bit outdated (pre-MGLRU): https://notes.valdikss.org.ru/linux-for-old-pc-from-2007/en/... Linux issue%3A poor performance under RAM shortage conditions
In Linux you get a ton of copy-on-write memory - every fork() (the most basic way of multiprocessing) creates a new process that shares all of its memory with parent. Only when something is written the child process actually gets "its" memory pages.
To put that into perspective, imaging you have only one process in your system, and it has a big 4GB buffer of rw memory allocated. So far so good. Then you fork() three times - your overall system memory usage is still roughly 4 GB. And now all four processes (parent and 3 children) overwrite that 4GB buffer to random values. Only at this point your system RAM usage spikes to 16GB.
This means, that the thing that actually OOMS may be just "buffer[i] = 1". It's very hard to recover from this situation gracefully, because this is an exceptional situation, and exceptional situations may require more allocations which are already impossible. Now compare that to Windows, where most memory allocations are in predictable moments, like when malloc() is called, and failures can be safely handled at that point.
So, in the ideal situation, Windows running out of memory will just stop giving new memory to processes and every malloc will fail. In Linux it's not an option, since every write to a memory location can suddenly cause allocation due to copy on write.
The rest of the system is still usable.
It's fine.
For the longest time I also ran with no swap on Windows (and just an excessive amount of memory). I'd notice when I'd run out of memory when a particularly hungry application like Affinity Photo died and I had a zillion browser tabs open, but again, the system is perfectly responsive and fine.
The Windows behavior seems much closer to deterministic and much more sane than the OOM killer of Linux.
Fortunately now we have MGLRU patchset, which "freezes" the active file cache for a desired amount of milliseconds, and in general is much smarter algo.
In a low-memory situation, the admin wants to ssh into the server and fix the problem that led into memory exhaustion in the first place. Whoops, MGLRU freezes the active file cache only, which includes the memory hog, but does not include sshd, bash, PAM, and other files that are normally unused when nobody is logged in, but become essential during an admin intervention. So, de facto, the admin still cannot login, and the server is effectively inaccessible. The only difference is that the production application is still responding, which is not so helpful for restarting it.
So out-of-the-box experience of some random distro is not necessarily the best you can get, especially on older kernels.
Are you referring to the /proc/pressure interface?
https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
I'm running Arch with i3wm. I didn't get ANY notification or helpful error message when I ran out! Instead, somehow, ghcup installed a corrupted version of cabal, that would segfault every time it was invoked. That was my only hint at first. I eventually ran df -h and discovered what was going on but man...
IME, if a process grows out of control, Linux won't notice until the whole system is thrashing, at which point it's too late and it tries killing random things like browser tabs way before the offending process.
In rare cases Linux might recover, but only because I hammered C-c the right place 30 minutes ago. In most cases a hard reboot is required (I left it overnight once, fans spinning at max, hoping the kernel would eventually accept my plea for help, but she had other priorities).
If you have enough memory that the desktop environment, browsers, and other processes that should keep running only use a small fraction of it, the OOM killer can pick a reasonable target reliably. A process that tries to allocate too much memory gets killed, and everything is robust and deterministic. I sometimes trigger OOM several times in an hour, for example when trying to find reasonable computational parameters for something.
About 8 years ago I got a work machine with 384GB of RAM, and I installed early_oom on it to make the OOM killer work a whole load earlier, otherwise the system would just become completely unresponsive for hours if one of my students/colleagues accidentally ran something that would make it run out of RAM.
I haven't really used any Windows version later than 2000 for anything except gaming, so I don't know how things work there these days. I mostly use macOS and Linux, and I've had far more trouble with pathological memory management behavior in macOS. Basically, macOS lets individual processes allocate and use far more memory than is physically available. When I'm running something with unpredictable memory requirements, I have to babysit the computer and kill the process manually if necessary, or the system may become slow and poorly responsive.
I'm guessing you are referring to "swapping", though?
If it's just one user process, it'll be killed by the OOM killer¹. That application will just be gone: poof. And for the rest you'll probably not notice anything, not even a hiccup in your Bluetooth headphones.
If it's many services, or services that are excempt from that killer, your system might start swapping. Which, indeed, leads to the case you describe.
¹https://unix.stackexchange.com/questions/153585/how-does-the...
It always felt strange that people buy lots of RAM but want it to be kept unused...
You could either get freaky with process explorer, or just keep some overhead so the system wouldn't try to do that. When I asked my guildies, they told me the 'default' for gaming is 16GB now, I was on 8 at the time.
Pretty much every gamer will at some point tab out to process manager to see wtf the computer is trying to do and exactly zero of them will think to themselves "I'm so glad there is no wasted memory!"
(edit: specified my intent on the reply)
Your approach is like buying a giant house, becoming a hoarder, and trying to throw a party.
exactly, excepting that the items they're hoarding are occasionally very useful for making their day to day activities go faster. And the hoarder has the superpower that in the blink of an eye they can discard everything that's hoarded to make room for the party.
Wait it isn't quite like a normal hoarder at all come to think of it!
Oh my god, it's 2023 and we're still discussing this idea from 1970s.
Is that so hard to grasp? No, stuff gets evicted from the cache long before you hit the swap, which is by the way measured by page swap-out/in rate and not by how much swap space is used, which is by itself a totally useless metric.
No...?
I'm looking at a machine right now that has 3.7GB not swapped out and 1.2GB swapped out. Meanwhile the page cache has 26GB in it.
Swapping can happen regardless of how big your page cache is. And how much you want swap to be used depends on use case. Sometimes you want to tune it up or down. In general the system will be good about dropping cache first, but it's not a guarantee.
> measured by page swap-out/in rate and not by how much swap space is used
Eh? I mean, the data got there somewhere. The rate was nonzero at some point despite a huge page cache.
And usually when you want to specifically talk about the swap-out/in rate being too high, the term is "thrashing".
if your cached disk pages keep getting hit and are "recent", they're going to stay in, and your old untouched working pages are going to be swapped out, to make room either for new working pages because you've just loaded new programs or data, or to make room for more disk pages to be cached because your page cache accesses are "busier" than some of your working pages.
you will swap out pages to make room for disk cache, but your cached disk pages will never be swapped out, they are just tossed (of course, after any dirty pages are written)
> Oh my god, it's 2023 and we're still discussing this idea from 1970s.
> Is that so hard to grasp? No, stuff gets evicted from the cache long before you hit the swap, which is by the way measured by page swap-out/in rate and not by how much swap space is used, which is by itself a totally useless metric.
Not everyone has been alive and into this stuff since the 1970s. That you and I know about this is irrelevant for the new people discovering it for the first time. There is always going to be a constant trickle in from new sources, for as long as it takes for the tech to go away. See relevant xkcd: https://xkcd.com/1053/
But it's also worth pointing out that RAM/swap/page cache isn't always as simple as page cache out, RAM in. For example, this question[1] seems to indicate that things aren't as simple as you suggest.
[1]: https://unix.stackexchange.com/questions/756990/why-does-my-...
As for the link you provided, I do think I can get a system in a state like that, and that isn't even untrivial. To push firefox into swap, esp if you have just 8 gigs of it is.. simple. But it is not in any way a normal state of a system. Idk how the author got it in that state.
Could've done away with less but I still have PTSD from all my applications crashing after I started Teams on my 16GB machine. On another note: Upgrading from i5-2500k to R7-5800X doesn't make Teams faster in any way.
It does allow you to save RAM and might prevent you from hitting swap for a while longer, but it won't save you if your working set is just too large and/or difficult to compress. Apps like web browsers with multiple tabs open might be easier to compress, a game with multiple different assets that are already in a variety of compressed formats, less so.
The Linux Kernel also has a bunch of optimizations (Kernel same-page merging, for example, among others) that do not require compression(although you could argue that same-page merging _is_ a form of compression).
The system is not supposed to 'lock up' when you run out of physical RAM. If it does, something is wrong. It might become slower as pages are flushed to disk but it shouldn't be terrible unless you are really constrained and thrashing. If the Kernel still can't allocate memory, you should expect the OOMKiller to start removing processes. It should not just 'lock up'. Something is wrong.
> which always takes minutes in my experience
It should not take minutes. Should happen really quickly once thresholds are reached and allocations are attempted. What is probably happening is that the system has not run out of memory just yet but it is very close and is busy thrashing the swap. If this is happening frequently you may need to adjust your settings (vm.overcommit, vm.admin_reserve_kbytes, etc). Or even deploy something like EarlyOOM (https://github.com/rfjakob/earlyoom). Or you might just need more RAM, honestly.
I have always found Linux to behave far more gracefully than Windows (OSX is debatable) in low memory conditions, and relatively easy to tune. Windows is a swapping psycho and there's little you can do. OSX mostly does the right thing, until it doesn't.
I don't why but locking up is my usual experience for Desktop Linux for many years and distros, and I remember seeing at least one article explaining why. The only real solution is calling the OOMKiller early either with a daemon or SysRq.
> It should not take minutes. Should happen really quickly once thresholds are reached and allocations are attempted. What is probably happening is that the system has not run out of memory just yet but it is very close and is busy thrashing the swap. If this is happening frequently you may need to adjust your settings (vm.overcommit, vm.admin_reserve_kbytes, etc). Or even deploy something like EarlyOOM (https://github.com/rfjakob/earlyoom). Or you might just need more RAM, honestly.
Yeah. Exactly. But as the thread says, why aren't those things set up automatically?
Has anyone transitioned from this being their observed behavior to something more tolerable? What did you change to avoid this problem?
I really wish this was a standard, configurable sysctl. There are many container environments (and heck, even browsers) that would benefit from this, and I cannot see any real downside.
I think with modern browsers, on memory constrained systems (think 4GB of RAM) this is easier to encounter than in the past. As someone who programs in Haskell from time to time I think I'm more familiar with Linux OOM behavior than most.
If someone wants to experience this easily with Haskell just run the following in ghci
foldl (+) 1 [1..]I've never seen a Linux system not lock up on OOM. At work or at home the instant it starts swapping you might as well restart. Of course it has to kill the SSH daemon rather than the process using 98% of the memory.
>I have always found Linux to behave far more gracefully than Windows
Windows just gets sluggish for a few seconds. You can even still move the cursor when that happens!
Unfortunately harder than it looks; if you compress the JS heap the garbage collector may decompress it again when scanning for references.
Also, Fedora has had zram enabled by default for a few years now along with systemd-oomd (which can sometimes be too aggressive at killing processes in its default configuration, but is configurable).
it's a lifesaver
I mean, swap is useful, but that's not what it's for. Same is true for compressed RAM. If you want an alert for low available RAM, it seems like it would be better to write a script for that.
Nope, neither of those things happen when zram starts compressing ram. Nothing grinds to a near halt until the compressed RAM space is used up, it just slows down a little bit. Btw, compressed RAM via zram isnt swap, it's available as actual ram. It also increases the total amount of ram available. I don't think I need to make arguments in favor of ram compression since Windows and macOS both have ram compression by default.
Wait, are you claiming RAM compression uses an adaptive compression factor that compresses more as memory pressure grows?
Are you sure that's how it works?
But yeah even with zram my laptop still hard reboots 80% of the time when it runs out of RAM. No idea how people expect the Linux Desktop to ever be popular when it can't even get a basic thing like not randomly rebooting your computer right.
I saw this also repeated at Apple (not Zswap since not Linux, but similar idea of compressing pages) and Android.
Edit: I should add, this is advice for desktops. If it's a server either resize or fix your service.
This is not true, there is a cost to freeing the cache pages and allocating them to the other program. I have seen some very regressive performance patterns around pages getting thrashed back and forth between programs and the page cache, especially in containers which are memory limited. You throw memory maps into the mix and things can get really bad really fast.
If you are interested in going deeper I recommend looking at the memory section of this book: https://www.brendangregg.com/blog/2020-07-15/systems-perform...
I have more then enough RAM on my office workstation to just accept this, but on my personal gaming computer that moonlights as a dev machine, I run into issues and have to kill WSL from time to time.
You might find this package helpful: https://github.com/arkane-systems/wsl-drop-cache
# sync; echo 3 > /proc/sys/vm/drop_caches2. Most things which report memory usage in a user-friendly way _already_ do this in an obvious way. (Htop shows disk cache in a bar graph, but doesn't add it to the "used" counter.)
3. Should UX always compensate for some fraction of users' misunderstanding of how their OS kernel works? Or would it be better for them to ask the question and then be educated by the answer?
Good UX makes the question "why is linux using my unused RAM for disk caching" (a non pressing question) instead of "why is linux eating up all my RAM" (panic, stressful question)
Yet on many Linux desktop you have to activate it (namely ZRAM). It solves the problem that a e.g. browser eats all your memory. It's much quicker than Swap and yet mostly unknown by many people who are running a Linux desktop. As mentioned by another user it's still not standard on Ubuntu desktop and I don't understand why.
It doesn't solve that. You get a little bit more headroom, but that's it. Not much ram is considered compressible anyway. On my Mac I'm barely reaching 10% of compressed memory anyway, so it doesn't make that much difference.
> In use: 18028 MB
> In use, compressed: 2718 MB
> Compressed memory stores an estimated 9013 MB of data, saving the system 6294 MB of memory.
That's not a small amount.
NAME ALGORITHM DISKSIZE DATA COMPR TOTAL STREAMS MOUNTPOINT
/dev/zram0 lzo-rle 15,6G 1,9G 248,6M 418M 16 [SWAP]What are you basing this on? Things in RAM are often very very compressible, usually between 3:1 and 4:1.
No offense, but you are being a very precise in defense but very broad in your (in general incorrect) claim.
The representation of an image sitting in memory will be a bitmap array, and for sure that will compress quite well. Video data as well but any decompressed frames are so transient I agree they won't benefit. ML-models don't compress well, but training data certainly does.
If you put aside mapping already compressed or non-compressable data into memory, all the rest of the things ram is used for can be compressed. Day to day you will have a lot of memory allocated that can be compressed. Most memory in use right now on most computers is compressible.
which I find easier to setup. Just enable it and it manages itself. You can still keep swap on disk, but it will act as a buffer in between, trading CPU cycles for potentially reduced swap I/O.
I think Arch has it enabled by default, but I am not sure about that. I had to enable it manually on Tumbleweed because my rolling install is years old.
If your web browser is using all your ram, it is probably misconfigured, maybe the ad-blocker has accidentally been turned off or something?
well it depends on your definition of modern, i suppose. i run Linux on a smartphone, which is about the most modern use of Linux i can think of, and hitting that 3-4 GB RAM limit is all too easy with anything touching the web, adblocker or not.
zram isn't exactly a trump card in that kind of environment, but it certainly makes the experience of saturating the RAM a lot nicer ("hm, this application's about half as responsive as it usually is. checks ram. oh, better close some apps/tabs i don't need." -- versus the default of the system locking for a full minute until the OOMkiller finishes reaping everything under the sun).
> There are no downsides, except for confusing newbies.
False. Populating the page cache involves lots of memory copies. It pays off if what's written is read back many times; otherwise, it's a net loss. It also costs cycles and memory to keep track of all these pages and maintain usage statistics so we know what page should be kept and which can be discarded. Unfortunately, Linux makes quantifying that cost hard, so it is not well understood.
> You can't disable disk caching. The only reason anyone ever wants to disable disk caching is because they think it takes memory away from their applications, which it doesn't!
People do want that, and they do turn it off. It's probably the number one thing database people do because they want domain-specific caching in userland and use O_DIRECT to bypass the kernel caches altogether. If you don't, you end up caching things twice, which is efficient/redundant.
Is data in the ARC double-cached by Linux's disk caching mentioned in the post? If so, is it possible to disable this double-caching somehow?
At least on FreeBSD, there is a kmem_cache_reap() that is called from the core kernel VM system's low memory handlers.
Looking at the linux code in openzfs, it looks like there is an "spl_kmem_cache_reap_now()" function. Maybe the problem is the kernel dev's anti-ZFS stance, and it can't be hooked into the right place (eg, the kernel's VM low memory handling code)?
(Bear in mind that 3 is the most aggressive but other than exporting the pool, it's the only way to dump the cache, especially if you boot off ZFS)
Newer versions of htop also now have counters for ARC usage (compressed or uncompressed)... but it still shows up as used rather than cache.
We observed that the real memory footprint for applications depends on many factors: file access pattern, disk IO speed (especially if swap is enabled), ssd vs hdd, application latency sensitivity, etc. Instead of coming up with some overly complicated heuristic, we use the Linux kernel provided memory.pressure [0] metric via cgroup v2. It measures the amount of time spent waiting for memory (page fault etc). Then by slowly reclaiming memory from the application until its memory pressure hits some target (say 0.1%), we can claim that the steady state usage is the actual memory footprint.
This may not be useful for PC but could be very useful for data center to track memory regression, and also to harvest disk swap without concerning too much about the cliff effect when the host runs out of memory and suddenly kernel pushes everything to swap space.
[0] https://facebookmicrosites.github.io/cgroup2/docs/pressure-m...
I'm running RKE2 on my desktop and it'll start killing pods due to low memory pressure, even though the memory was only used for disk caching. I wonder if there is any way to make it stop doing that and instead only start killing pods if it's due to "real" low memory pressure.
[0]: https://github.com/kubernetes/kubernetes/issues/43916
[1]: https://docs.kernel.org/admin-guide/cgroup-v2.html#memory-in...
[2]: https://biriukov.dev/docs/page-cache/6-cgroup-v2-and-page-ca...
[3]: https://www.freedesktop.org/software/systemd/man/systemd.res...
And the response was similar: all that "used" RAM can be reclaimed at any time should an app need some, but in the meantime, the system (which is Linux) might as well use it.
I think they "fixed" it in later versions. I don't know how, but I suspect they just changed the UI to stop people from complaining and downloading counterproductive apps.
As usual in these situations, unless you really know what you are doing, let the system do its job, some of the best engineers with good knowledge of the internals have worked on it, you won't do better by looking at a single number and downloading random apps. For RAM in particular, because of the way virtual memory works, it is hard to get an idea of what is happening. There are caches, shared memory, mapped files, in-app allocators, etc...
This has increased the quality of our memory monitoring by a lot.
Whether swap is available is more or less irrelevant for this behavior. The only thing that swap changes is that kernel is then able to “clean” dirty anonymous pages by writing them out to swap.
Fwiw without swap there isn't really any paging in or out (yes mmapped files technically still can but they are basically a special cased type of swap) so your question is hard to parse in this context. The disk cache is all about using unallocated memory and an allocation will reduce it. Paging is irrelevant here.
Btw you should always enable swap. Without it you force all unused but allocated memory to live on physical RAM. Why would you want to do this? There's absolutely no benchmarks that show better performance with no swap. In fact it's almost always the opposite. Add some swap. Enjoy the performance boost!
https://haydenjames.io/linux-performance-almost-always-add-s...
Whatever conceivable speedup there is from 12 GB of file cache as opposed to 11 is obliterated multiple times over from the time lost by having to do this dance, or worse, recovering after the oom killer wipes out my X session or browser.
Perhaps you can share more details of what you're doing to force the cache to drop and what the side effects are exactly because an OOM can't be caused by the file cache since the total free memory available to applications remains the same. The whole point of the file cache is to use otherwise unallocated memory and give it up the moment it's needed. There should not be an OOM from this short of an OS bug or an over allocated virtualized system.
echo 3 > /proc/sys/vm/drop_caches
Last time I ran into this was a couple of years ago on a stock Arch system. (Disabilities forced me back to Windows). Every time, the largest memory consumer was the web browser. Also every time, the system became nearly unresponsive due to swap thrashing (kswapd at the top of the CPU usage list, most of which was I/O wait).Last time I complained about this problem, someone suggested installing zram which did stop it from happening. However, this does not change the fact that there is some pathological failure case that contradicts the central thesis (not to mention, smug tone) of this website and makes searching for solutions to the problem infuriating.
I was too lazy to find a proper solution, so I just used mlockall after allocating a massive heap and pin the process to a core that is reserved only for this specific purpose.
I think cgroups has very flexible tools for reserving system wide resources, but haven't had the time to test it yet.
I build older android (the OS) versions inside docker containers because they have dependencies on older glibc versions.
This is a memory-heavy multi-threaded process and the OOM killer will kill build threads, making my build fail. However, there is plenty of available (but not free) memory in the docker host, but apparently not available in the container. If I drop caches on the host periodically, the build generally succeeds.
2. IME the kernel takes the container's memory limit into account when determining whether to allocate a page for cache. Caching, by itself, won't cause the container to exceed a memory limit.
This is a reference to a legitimate piece of Internet history:
https://en.wikipedia.org/wiki/Ate_my_balls
> "Ate my balls" is one of the earliest examples of an internet meme. It was widely shared in the late 1990s when adherents created web pages to depict a particular celebrity, fictional character, or other subject's zeal for eating testicles. Often, the site would consist of a humorous fictitious story or comic featuring edited photos about the titular individual; the photo editing was often crude and featured the character next to comic-book style speech in a thought balloon.
> The fad was started in 1996 by Nehal Patel, a student at University of Illinois at Urbana-Champaign with a "Mr. T Ate My Balls" web page.
'man free' states
used Used or unavailable memory (calculated as total - available)When you can download as much RAM as you want, any time you want.
No downsides except for massive data loss when the system suddenly loses power/a drive crashes and the massive theft of memory from host OSes (e.g., when using Windows Subsystem for Linux).
I don't know enough of Linux internals to know if the writeback cache and read cache are the same object in the kernel, but they feel similar.
Of course the real response is that without write cache (effectively adding fsync to every write) any modern linux system will grind to absolute halt and doing anything would be a challenge. So contray to GP's post, it's not reasonable to complain about it's existence.