I've tried using a VM with overcommit turned off and a modest amount of memory. Among other things, my mail reader, mutt, used more than half the system's memory when looking at my mail archive, so it couldn't fork to exec an editor to write a new mail.
It doesn't have to wait until it actually runs out before issuing ENOMEM. Most processes don't deserve to starve the system completely.
Has this application already allocated 90% of the memory? Has it been steadily growing the allocation without releasing much back? Well then, why let the system run out completely? Why not stop it at say 10% or 5% left?
For a server, perhaps, but I'd say even for a server it would be better that say the core application returns a 500 internal error or whatever than forcing the system to start killing random processes.
I don't know how common this is on the Linux/POSIX side, but in DOS and Windows it is usual practice to allocate some amount of memory at startup and use that for error-handling code, so that things like showing error messages will not cause any more allocations.
It's also hard to get a good idea of how much memory a program is really using, which makes setting reasonable limits for things tricky.
The best you can do is to have a small swap space, and alert when it gets to 50% and fix whatever. But then you have the problem of filesystem pages evicting anonymous pages to swap which ruins the utility of the swap usage as a gauge and/or drives you to configure way more swap than is reasonable. Although, maybe a bigger swap space and alerts on swap i/o rate might be ok, other than it's awful painful to have swap space of even 0.5x ram if you have 1TB of ram.
The middle ground is your own config unfortunately - cgroups can limit available memory, but you'll have to set it up by hand.
You cannot make reliable systems thinking like that. The point of resilience is __not__ to rely on correctness of kernel's advertised behavior nor correctness of your assumptions about it.
One cannot build highly available systems and networks on unreliable, lying software. Basic guarantees are required. If one has 10,000 systems and one loses power on even a portion of such lying systems without basic guarantees, no amount of distribution will guarantee data consistency. I'm no stranger to building very large distributed clusters, but those builds start with an operating system and software which do not overcommit and which are paranoid about data integrity and correctness of operation. In fact, I'm specialized in designing such networks and systems, from hardware to storage all the way up to application software.
Many databases rely on the overcommit being possible.
For the general case though, because of the way programs have been written, it is easier to have overcommit on.
Linux is completely useless with ram that is almost full in a way that OSX and windows absolutely are not.
One way to make relatively sure that this is always the case is to use a userspace daemon like earlyoom.
Though to stay fair all desktop OS'es behave badly when put under memory pressure, it's just that Linux is an order of magnitude or so worse.
On the other hand, syscalls across architectures of the same bitness are exactly the same, minus the newer architectures just not having some historical abominations like sbrk. (IIRC, syscalls are different between even aarch64 and amd64 on Linux, which is just… why?!?)
We are stuck with needing to support x86 and prefer to do so as seamlessly as possible for the end user (call us old school like that); to that end we have a single bootable ISO image containing our appliance that needs to run on generic x86 hardware ranging from cutting edge to 15+ years old (I've posted about that with tons of accolades plus patches maximizing compatibility to the FreeBSD dev ML [some|a long] time ago but never heard back). We ported the software over from a very messy Linux-based implementation to a much cleaner architecture under FreeBSD 9 what I now realize can be considered a pretty long time ago.
With that background out of the way, we had written a custom scfb-powered "guaranteed compatible" X display driver (back before there was an upstream scfb driver, although I am sorry to see even with FreeBSD 12.0 that xf86-video-scfb does not work on quite a number of hardware profiles) that would work so long as the framebuffer did (and it usually does). I remember running into an issue (under 10.0) where even basic FBIO_* ioctl defines had different values when compiled for AMD64 vs i686, so although our x86 Xorg module would run under x64 (with the i386 compatibility module loaded, iirc?), it would fail to work as the wrong driver routine was being called into. I was honestly surprised it didn't work oob, as these were very basic ioctl calls.
Of course today with UEFI, signed drivers, and everything else, it's a miracle we haven't already been forced to distribute separate ISOs, but that's really only because we gave up on signed booting and simply require a legacy CSM for UEFI hardware, which is more steps but generally universally available; I don't think there's any real impetus behind getting UEFI support for x86, so we still boot the old-fashioned way.
Getting completely off topic, honestly both Linux and FreeBSD are quite literally decades behind when it comes to being able to guarantee unaccelerated video output at {,e,s,x}vga resolutions without requiring hardware-specific drivers; I don't think I have ever seen a MS-DOS/Windows PC that wouldn't at least boot in VGA mode going back at least all the way to early 90s. We certainly spent no time agonizing over getting video to work when Microsoft used to allow us to license WinPE to power our appliance images. I have an unfinished patch somewhere for the vt VGA driver that adds support for the various ioctls, but it seems the world has unfortunately decided that going with hardware-specific (k)drm drivers is the way forward. We decided to bite the bullet and now ship with those too, but it's about the farthest thing from bulletproof compatibility and neither the vendors nor the reverse-engineered community driver devs really bother formally testing these modules and it unfortunately very much shows. And now nVidia has discontinued providing proprietary x86 drivers for FreeBSD and Linux (although they notably still do for Solaris, almost certainly because of the aforementioned ABI compatibility guarantees there), so we're probably going to need to split our legacy/i686 and the uefi/x64 images soon enough.
> signed booting
It's not like FreeBSD was an early adopter of that. Some kind of secure boot support things have landed in the last couple months, and veriexec for signed stuff, but very few people have started using this.
> separate ISOs
Or a "fat" ISO that has both 32 and 64 bit versions?
> nVidia proprietary drivers
With these, there's already an issue with which versions of that support which cards..
QEMU supports user-space emulation on Linux, for example you can run an ARM binary on top of a amd64 kernel. x86 on top of amd64 kernels works OOTB.
Presumably it must also include some sort of compatibility layer rather than merely translating machine instructions, as the syscall payloads vary considerably. But qemu is certainly small and (when used correctly) fast enough that depending on that should not be a dealbreaker if you really have no better options.
Anyway Linux won for now, but Solaris lives on vicariously through illumos and SmartOS, which handle this the same way as Solaris does (it's the same code) and the codebase is very actively developed, so it's not over yet no matter how much Linux is winning now: I'm waiting for the folks to figure out just how shitty Linux is under the hood and the fact this is a topic on HN tells me it's starting to happen as people gain more and more experience. Linux won the battle but with such shitty code underneath winning the war as history will judge it is another matter entirely. Meanwhile SmartOS continues to be developed and works correctly in these situations.
Prior to FreeBSD 10, I'm paraphrasing but essentially it would pretty much not swap anonymous pages until memory pressure got really high -- it would evict clean disk pages and try to clean dirty disk pages first. But if the pressure continued, it would start marking anonymous pages and eventually swap those out too. A big problem was there was a massive slowdown when it hit the threshold; tons of pages would be marked as inactive, so using them in the process triggers a page fault (and they then get marked active), and all the page faults and page table activity slows things way down.
In FreeBSD 10, a patch from NetApp changed the page scanning behavior so it was always running. This is good, because you don't get that massive slowdown while the pages are determined to be active or inactive; instead the kernel always has a much better idea of which pages are used; but it's also not great because some of the inactive anonymous pages are going to get paged out, and it also makes memory accounting trickier. In FreeBSD 9, you could look at the active memory stat on a system that wasn't at the brink and figure that was pretty much the amount of ram you needed -- so if it was growing over time, you might have a memory leak, or real growth. In FreeBSD 10, it's harder, because your programs may need things loaded but not access them for some time, and the pages could get marked inactive, so you really have to guess; and swap usage isn't a great metric either because those idle pages might get swapped out. You can set the vm.defer_swapspace_pageouts sysctl to get something similar to old swap behavior, but you still lost the memory usage gauge.
OOM behavior seems better, but also not necessarily great. Most of my experience is from running Erlang servers with very little other stuff going on, so it might not apply to a more conventional server with lots of random things, and probably not to a desktop either. Overcommit is disabled by default (and i think it may be substantially different than Linux overcommit anyway), and mostly what happens in my experience is the big allocator (beam.smp for us) gets its allocations failed when you run out of swap and then it crashes and everything is in an ok state. Every once in a while, something else will get failed allocates and die before the big allocator, but then the big allocator still dies; this wasn't great because sometimes it would be an important but small daemon like ntpd, when this happened we'd often reboot to get back to a known good state. There's also an OOM killer that can be triggered depending on exactly how things go down -- it just kills the biggest process; for us, that's the right thing, if the VM is eating all the ram, killing off gettys isn't going to do any good.
However --- we saw several times that the kernel would just hang. Unfortunately, when we were seeing this often, I hadn't figured out the kernel debugger, and it was usually in the middle of an incident anyway, so we'd reboot and go on with life. A hung kernel is better than a thrashing kernel, but only a little bit.
No way. Processes allocating terabytes (GHC Haskell compiled programs, AddressSanitizer, etc) work fine out of the box.
Reading tuning(7):
> Setting bit 0 of the vm.overcommit sysctl causes the virtual memory system to return failure to the process when allocation of memory causes vm.swap_reserved to exceed vm.swap_total. Bit 1 of the sysctl enforces RLIMIT_SWAP limit (see getrlimit(2)). Root is exempt from this limit. Bit 2 allows to count most of the physical memory as allocatable, except wired and free reserved pages (accounted by vm.stats.vm.v_free_target and vm.stats.vm.v_wire_count sysctls, respectively).
Soooo vm.overcommit=0 does not mean "no overcommit" (which would make sense), no, it seems to mean something like "no special flags like disabling overcommit" :D
Perhaps the real bug is that Linux distros make it easy to run swapless.
Before quitting background applications it first sends them a request to free memory, in a well-behaved iOS program you use this to clean up your caches and ensure your don't use more RAM than you absolutely need. You should also suspend your state to disk when your app is backgrounded so you can just continue where you left off if your app is killed.
Many macOS apps also do this, you can forcefully restart a Mac and after a reboot it'll restore your session to pretty much the exact state you left it in, including any open 'unsaved' files.
Linux could implement a similar mechanism to signal apps to clean themselves up and maybe a 'save your state, you're about to get killed' signal.
Isn't that pretty much what "memory.pressure_level" [1] is?
[1]: https://www.kernel.org/doc/Documentation/cgroup-v1/memory.tx...
Almost all of them get closed because they get old and no-one is willing to do the necessary work to refine the bug report to dependable reproducibility.
And that's not surprising really; it's hard, time-consuming work which may be obsoleted on the next kernel version.
That said, some have written memory stressors and have been able to crash or stall a machine but it's still a bit hit and miss.
A better one would just make a C program that mallocs more memory than is available. The "open enough tabs so that it crashes part" is like "banana for scale", it is incredible unspecific. You could probably open HackerNews 50x more than espn.com