I've now disabled systemd-oomd on my Fedora desktops
utcc.utoronto.ca
utcc.utoronto.ca
I think fork() tends to be implemented using COW memory. Would it help to deprecate fork+exec in favour of posix_spawn()?
Alternatively, just let programs opt-out of overcommit and also OOM killer. Then only badly behaving programs will be shot in the head randomly.
If there are some situations where there is no logical course of action? Then yeah, crash I guess. Assuming your software isn't responsible for people's lives. It's still better than the OS randomly killing you because some other program fucked up.
I mean this is just crazy. The OS lies to you about memory, you ignore it anyway since it always lies to you, and then we randomly kill programs when we're out of memory. This is not a reasonable way to do things.
The problem, as I see it, is that by now programmers are so used to not paying attention to memory (or any other resource in most cases) that there's a ton of inertia behind the "nothing can be done" sentiment.
$ sysctl vm.overcommit_memory=2 # policy number 2
$ sysctl vm.overcommit_ratio=0 # ratio = 0%
So here’s your golden opportunity to show us how easy it is to program against that!I didn’t claim that it is not easy. I claimed that it is practically impossible. You claimed it was easy. So please, prove me wrong and just show us how easy it is.
> I claimed that it is practically impossible.
Suddenly it makes sense why modern software is such unreliable slow garbage. You don't by chance work for Microsoft on the Windows Terminal team do you? If this is so impossible, why would so many other OSs actually correctly not-lie to the program about memory allocation failure?
> So here’s your golden opportunity to show us how easy it is to program against that!
Alright, what would you like me to program for you for free and how long are you going to give me[0], because the vast majority of code I write doesn't even have to allocate at all. Even the stuff that does just does the obvious thing when it can't allocate what it needs: tell the user it can't allocate what it needs and exit as gracefully as possible. Even that is impossible on an OS that fucking lies to you.
[0] No, I'm not going to actually do it. This whole tactic is predicated on the idea that it is more work than someone is willing to go through to prove a point on the internet, you know that, that's why you went there.
You don’t need to go and write software for me, if your arguments are worth anything you would be able to demonstrate them using the software you already wrote as you gained the experience to back your handwaving claims. Clearly you didn’t do that (if you did you would have known your original claims were a bit shallow), you just show up on your high horse to talk down to others who do. All talk and no action.
Which is ridiculous, because plenty of programs will stop and tell you they have run out of memory for an operation. It's the easy way to handle allocation failure you're so hell bent on me showing you.
Here's one: https://knowledge.ni.com/servlet/rtaImage?eid=ka03q0000012W2...
You can do that kind of thing when the OS doesn't lie to you. Depending on circumstances you can do other things.
But hell, even if you think that's stupid and I'm stupid for saying that this is better than crashing, I still can't fathom why you think the status quo on Linux is the best possible idea. Hey, I know, let's pretend we can't run out of disk space either and just randomly delete files when the disk starts to fill up.
If you do one huge allocation and it fails, yes it’s easy to handle that. But programs do tons of small allocations all the time and it’s impossible to handle if that fails. Do you think you can even show that fancy pop-up when you can’t allocate memory?
And the OS behavior has everything to do with it, that's the whole context of this discussion: Linux lying to programs about memory and having to deal with the consequences in a really dumb way.
> But programs do tons of small allocations all the time
Mine don't, at least they don't allocate a bunch of small shit from the OS all the time[0], that's slow. Again I suspect learned helplessness here: people have been doing it a certain way for so long they don't even consider alternatives.
[0] Actually, come to think of it I don't think most malloc implementations do it differently.
This is different from whether applications handle this reality well (by and large, they do not). You are 100% correct with this. As is the poster who says fixing it is next to impossible (given the wide amount of deployed, critical software). I think if you actually dive into this, you will find Windows does not handle this as gracefully as you want to believe it does.
Everyone says "turn of overcommit" as if it is a general solution. This works only when you have tight control of what runs on the system, and it is tailored for a specific application or deployment that behaves well. And guess what? If this is you, then you already know how to disable overcommit in the kernel build for the custom OS you are building for your product. I will leave it as an exercise for the reader to determine if the policy of crashing the current process is a better strategy than what the OOM killer does (it favors to reap shorter-lived and larger-RSS processes).
Linux is not Windows. People get upset when they first realize this. They want different defaults (nevermind that they can customize it however they want to, and it isn't as though this is some esoteric kernel topic buried in the lkml that should surprise anyone).
If anything, I'd say that Linux needs easier ways to securely load a different set of code for running at OOM time.
Right, but it is only necessary because Linux lies to the application. If Linux instead reserved memory for itself to remain responsive and properly allowed allocations to fail when they would exceed the available resources, the applications that ask for memory that is unavailable would be the ones having to deal with the consequences rather than the OS having to kill processes based on some heuristic.
> I think if you actually dive into this, you will find Windows does not handle this as gracefully as you want to believe it does.
Personally, when I've seen modern Windows systems grind to a halt it is because of disk IO saturation, not memory saturation. Maybe I'm just not in the right contexts to see it.
> You are 100% correct with this. As is the poster who says fixing it is next to impossible (given the wide amount of deployed, critical software).
Ok, so that may be speaking at cross purposes then. I mean on an individual application level this is a not-too-hard solvable problem (usually). If they mean "given the current state of the ecosystem, it is globally nearly impossible" then I can see where they are coming from, but again I have to wonder what the point of having all this software be FOSS is if we can't work towards a goal like this.
> Linux is not Windows.
Yes, and in many ways I find this unfortunate. Good ideas are good ideas regardless of which OS they come from, fanboism is stupid.
> They want different defaults
Because defaults matter. As mentioned, because of the default applications have been written a certain way over years and years and now it is not a simple matter of disabling overcommit because of what happens to that software when you do that. That's why I proposed a solution based on opt-in.
I guess my entire point with respect to defaults is that for anyone who knows what they're doing and the defaults do not work for them, they already have a huge pile of systems design and architecture work to do, and the mechanics of changing this OOM policy for that is a trivial change.
For people who do not know what they're doing and want the OS to do the hard stuff for them, there is no sane default. There are tradeoffs that will make large numbers of people unhappy. It is a no-win situation.
I don't understand why you believe that forcing the process that is requesting the allocation to deal with it is the best thing generally. From a systems perspective, it is actually worse because there's a stronger chance you're killing a nice, stable process that might be critical simply because it lost the lottery.
The OOM killer strategy is to try to reap short lived processes that allocated a lot. From a systems perspective, this is clearly better... though as I said it might be better if you could more easily modify the OOM heuristic.
Because it is in the best position to understand the consequences of failure and the options for dealing with it.
> From a systems perspective, it is actually worse because there's a stronger chance you're killing a nice, stable process that might be critical simply because it lost the lottery.
If the process is truly critical, it should be designed to stay stable in the event it can't allocate memory.
One could at least fail gracefully: at start-up, request enough memory for a buffer reading ‘request to allocate memory failed; exiting’ and then print it to standard output, then exit.
One might initiate a garbage collection cycle.
One might even allocate enough buffers at start-up to cleanly process outstanding work, write results and then terminate.
It requires a little thought and engineering, but there are approaches which don’t just dump core.
You can actually take some ideas from the Linux kernel if you like because it has what you propose, a way to allocate memory that never fails and never blocks. And there you can tell how easy that is.
So, all of them then?
I speculate it isn't common, or at least it isn't required for ordinary desktop apps. Windows survives without fork().
> And isn't posix_spawn implemented in terms of vfork, which a lot of people like even less?
That sounds like a glibc implementation detail, not something required by the standard. I'm not an expert however. Simply, again I come back to: Windows can spawn processes without using it as an excuse to overcommit.
From a science perspective, the ability to fork is really convenient for parallel processing. Load some data and then fork a pool of worker processes, and they can all read almost free copies of that data. This is much easier than setting up shared memory. It's a bigger deal on an HPC machine with 128 cores than a laptop with 8 cores, but even on a laptop it's a significant point in favour of Linux.
Getting rid of it in Linux would also be going against Linus' rule that the kernel never breaks userspace.
https://www.kernel.org/doc/Documentation/vm/overcommit-accou...
But an app which thinks it has successfully been allocated memory, being chosen at random to crash? Not ideal, just like in OP's example.
Doing this wont give you the fix you think it will.
It's not systemd-oomd that is the main problem here, it's Fedora's implementation/application of it to all the user@.service. systemd-oomd does suck at logging, but that seems to me like a silly deal-breaker to throw out the baby with the bathwater. It can (and will) be fixed.
We aren't. Our time's up. The new generation is on now, and we didn't teach them hardly anything we learned because we were so invested in making fun of them for enjoying ubuntu.
Best way we can move past it is to take the next curmudgeonly step and move to OpenBSD and Plan9, as they're the only things sufficiently dense and opaque to be as yet out of reach of the current generation.
OOMD is a workaround for this behaviour.
And it can't be disabled. I disabled the kernel memory overcommit and then neither Chrome nor Edge were running.
I can't use my computer without my browser of choice. Therefore, overcommit can't be disabled.
The browsers will allocate more memory than is physically available in the system but often not "touch" it. Without over-commit this causes an OOM situation and the kernel kills the browser, but that's what it's supposed to do when over-commit is disabled!
I believe some of these massive allocations are used for security purposes or somehow related to the chromium sandbox but I can't remember exactly what's going on. Some sanitizers like ASAN do something similar for detecting invalid memory accesses (and hence can't be used without over-commit either)!
I thought systemd is open source and if you don't like something about it, you can just submit a patch or even fork it or just not use it.
Replacing system components is hard and the early days are going to suck. Just because there is a lot of pain and suffering with Wayland does not mean that it was a mistake and we should have just stuck with X11.
Computing has changed since 1970, some things in our unix-like systems will also need to change to accommodate the current world and some of it is going to be painful, that's just how it is, we made fun of Vista, now it's our turn.
If the default configuration were good and you changed it to be bad, that's entirely your fault. But if the default configuration were bad and you just didn't fix it, that's mostly if not entirely upstream's fault.
> I thought systemd is open source and if you don't like something about it, you can just submit a patch or even fork it or just not use it.
Yes, we can do that. The point is that we shouldn't have to.
> Replacing system components is hard and the early days are going to suck.
Nobody's complaining that things are getting replaced. We're complaining that the replacement is worse than the old thing.
> Computing has changed since 1970, some things in our unix-like systems will also need to change to accommodate the current world and some of it is going to be painful, that's just how it is, we made fun of Vista, now it's our turn.
Systemd didn't exist until 2010, so any changes from 1970 to 2010 clearly can't have made it necessary. And the BSDs are all still good, functional operating systems, despite not being able to run it, so it's clearly not necessary even for what's changed since 2010.
My desktops have 64GB+ of memory typically, they're virtualization workhorses. I don't know what it is in oomd, but any time I make actual real use of it, things started getting killed.
Zero memory pressure. Talking sections of 60+GB free, so it's not fragmentation either.
The grouping of the entire session the link suggests seems likely
THE Pain point remains in systemd-oomd's poor (and often, no) crash dump or traceback support.
Nevertheless, systemd-oomd remains an excellent OOM mass-killer (via cgroupv2), just not the discriminating assassin ... as (mis-)configured ... by Redhat/Fedora.
When I've seen them complain the systemd people tell them that they're just stuck in their ways and that systemd is a revelation on desktops where users don't have the historical unix experience holding them back...
I've had critical systems fail on reboot because the NICs swapped names so WAN port got configured as LAN and vice versa. That can also create a terrifying security problem too
Which also was wrong when switching hardware around, but that's a different issue.
eth0 tells you the first ethernet class interface that was activated that boot. Not particularly helpful.
I feel like if you need to care about these names as a user... the distribution has failed to provide effective wrappers. This is window dressing.
While sure, I would like shorter device names, the fact is... the only time I need to use them is when ambiguity is something I/we can't really afford.
So, I'm thankful for the cryptic-yet-specific names.
Everyone else (read: desktop people) can use whatever DHCP'd.
At least it doesn't require a reboot. Yet.
Other times, like oomd, I just want it to get out of the way.
Service dependencies are an area I love. With systemd making a service robust around dependencies like mounts and other services is a cake walk.
If you have manual hand-holding processes on services running under systemd, I suggest looking into supplying these kinds of directives.
- Requires / Wants
- Before / After
- PartOf
I'm torn for networkd. The configuration of interfaces is nice and simple, I've grown to like the link, netdev, network concepts/files.However, forwarding is significantly different. I'd think twice about using it on a system doing NAT/masq
Init scripts suck to test and debug. When manually-called, the caller's environment would get used, rather than a clean env, which would cause problems with the poorer-quality scripts. These showed up multiple times in production where an engineer would bounce a process manually and it got some configuration from their environment. We took some steps to alleviate this, like encouraging use of `service`, but using the scripts directly was reflexive for most because of years of indoctrination and it's portability to some other operating systems. systemd eliminated this entire class of issues, because unit files are just config files, not scripts.
Dependencies actually work. Usually, this came up when needing to wait on the network. With systemd, this is just a single dependency in the unit file. It also meant we stopped having to manage S## and K## which became something of a time-consuming art on machines with lots of services.
Not exactly init system, but systemd's triggered units absolutely destroy cron. Manually running a job (e.g. when testing or when a run failed) guarantees the same environment, whereas cron has no similar capabilities. It also actually keeping status means we can more easily see if the last run succeeded. This may ease our monitoring burden, too, but we haven't gotten to give it attention yet.
It’s only useful on desktop because of its similarities with launchd, which is particularly good on desktop.
On the other hand, not having a userspace OOM killer means (meant?) that if your system runs out of memory, it will become so slow as to be totally unresponsive and the only thing you can do is to reboot. Which is even worse. Is that still true today or did the kernel OOM killer improved?
How does earlyoom compare against systemd-oomd?
Or, as a mnemonic: Reboot Even If System Utterly Borked*
*Being Irish, I prefer "Banjaxed" for the last word, personally.
You got me to look it up: Alt-SysRq-F just for the kernel oom_killer ...
On ThinkPads, that's: Press Fn-Alt-S, then release Fn-S, then press the desired letter. Result from dmesg: "[154256.637037] sysrq: This sysrq operation is disabled." Bleh. Didn't work when you needed it :) Of course these can be enabled, but earlyoom is soo much easier than remembering the OOM dance or the REISUB dance (particularly if it is just an OOM :D)
It's been well over a decade since I've seen that kind of behavior. Storage devices became much lower latency and the kernel swapper got much smarter.
Linux Desktop in OOM situations is absolutely trash.
You're not suckless. Nobody wants you to be suckless. I like you specifically because you aren't anything like suckless.
Don't replace anything unless you have close to full feature parity. Nobody wants your minimum viable product components. The people who do don't want systemd.
Init system? Absolutely. Timers? You've done a wonderful job. Time sync? Please no, crony does it better.
> Achievement unlocked: I just had systemd-oomd kill my entire X session for unclear reasons.
> Does it log details about the session it killed so I could see, say, which program was using too much memory?
> Nope.
>
> Systemd-oomd has now been voted off my island.
https://github.com/systemd/systemd/blob/v252/src/oom/oomd-ma...https://github.com/systemd/systemd/blob/v252/src/oom/oomd-ma...
Chris's wiki has now been voted out of my sight.
Sometimes you wish X had been killed, instead of a total freeze or for example LUKS/dm-crypt being unable to allocate memory (and causing way way way worse issues than X being killed).
Yeah, looks like logging was added 1.5 years ago so hard to take his perspective seriously.
Considering systemd-oomd is a relatively new addition to the project, contributed by FB staff, it wouldn't surprise me if there were desktop integration issues yet to be sorted out by the various distros.
The post doesn't strike me as made in good faith, seems written by someone with an axe to grind with systemd.
The most facepalm one was somebody uploaded their `.vimrc` which had a bug that infinitely recursed the parent directory loading ctags until it gobbled up all available memory and crashed. So every time he opened vim, it started about a 40 second timer of doom unless vim was exited before memory exhausted. Took forever to figure it out too, because he often only opened files for 10 seconds or less, short enough not to trigger the OOM.
systemd-oomd (and earlyoom, which Fedora used for a couple of releases before switching) try to detect that you're running out of RAM before the system becomes unusable, and kill a big memory user. It periodically killed Firefox before I bought some more RAM, which was a pain, but less of a pain than hard-rebooting the laptop.
The plans for Fedora to do this are public pages:
https://fedoraproject.org/wiki/Changes/EnableEarlyoom https://fedoraproject.org/wiki/Changes/EnableSystemdOomd
Ubuntu 22.04 was also affected: https://news.ycombinator.com/item?id=31530739
And flipped a config flag to make it less aggressive: https://www.phoronix.com/news/Ubuntu-Drops-Swap-Kill
It also gives me a nice desktop notification about what it killed, though I'm not sure where/if a more permanent log is written.
Cost me god knows how much money or wasted power from lost computations until I figured that out and disabled it.
This might have made some IaaS firms happy?
From the point of view of the program the files are written under /tmp, but in reality they're stored somewhere else. Try this to see the files:
ls /tmp/systemd-private-$(systemd-id128 boot-id)-httpd.service-*/tmp
Run systemctl show --all httpd.service
and look out for the PrivateTmp directive [1]. Use systemctl edit httpd.service
to override it.[1]: https://www.freedesktop.org/software/systemd/man/systemd.exe...
1 I use std function to get a temp file path
2 I write content in that file
3 I do all the checks to ensure the write worked, I use std functions to check file exists
4 I use std function to print the file path
And the file path is a lie, It all would be fine if it would use the real path, and if a newb would hardocde "/tmp/file1" then I would be happy with a error message that I can't write there .
I assume they had some reasons for the lie, I am not convinced it was worth it, I do not like magic hidden stuff.
The "antifeature" can be disabled if you don't like it, but keep in mind that it increases security. There have been vulnerabilities related to temporary files [1][2]. Usually there is no need for a web server running in production to make its temporary files accessible to the rest of the world.
The "lie" is similar to what happens when you're using virtual machines. The virtual machine thinks it's using some hardware, but in reality...
On a side note, Flatpak also does this by default sometimes. It can run applications in a sandbox in which even the home directory is partially inaccessible.
[1]: https://cwe.mitre.org/data/definitions/377.html
[2]: https://nvd.nist.gov/vuln/search/results?form_type=Advanced&...
[Service]
PrivateTmp=noBut it's such a misfit on desktop distros that the criticism there is valid and much louder. It never should've been a one-daemon-fits-all solution.
This was surprising, I was unable to replicate it though. Did you set systemd to automatically clean up sessions?
From the man page of logind.conf,
KillUserProcesses= Takes a boolean argument. Configures whether the processes of a user should be killed when the user logs out. If true, the scope unit corresponding to the session and all processes inside that scope will be terminated.Plus, on a theoretical level that is the correct behavior, you don’t want to leave running processes gobbling up resources under a logged out user without explicitly allowing it.
Correct behavior is that a demonized process doesn't ever get killed just because another process in its original family tree dies.
Who decides what is a demonized process, vs one that just froze and won’t respond to some random UNIX convention, that can’t even reliably signal its state (what is tmux when it is frozen?)
So the solution is to explicitly specify that a given process wishes to live outside of the seat’s lifetime (either by specifying it as a service, which your earlier example ssh is on basically any distro, or by executing it as with systemd-run)
Okay, so it does that now. But now how do you tell if it's frozen?
The point is that the OS does actually know “statically” that a given process is a daemon or not.
The solution of course is more systemd. Instead of using generic Linux tools like grep, learn and remember a bunch of weird systemd-journald specific commands. Unfortunately I'm also a programmer, and only really have room in my head for one grep command, so it introduces frission and I have to look up the commands every time.
Also just a lot of random breakage over the years.
Service management is a hard problem when you start adding things like socket activation, home directory management, DNS resolving, container management, log management, and whatever else they want to throw in.
Now you might say "well DNS resolving isn't actually part of systemd, it's a seperate daemon", well they've been more and more tightly coupling all these different components together. Yes you can disable systemd's DNS resolver, but probably something else will break. They're all pretty tightly coupled together and they mostly seem to do that in a dumb way.
I can also see the design decision made to kill the graphical session when the system becomes unresponsive. The alternative to that situation pre-systemd-oomd was a force reboot of the system.
This is so obviously poorly thought-out as to border malicious negligence.