I've tried using a VM with overcommit turned off and a modest amount of memory. Among other things, my mail reader, mutt, used more than half the system's memory when looking at my mail archive, so it couldn't fork to exec an editor to write a new mail.
It doesn't have to wait until it actually runs out before issuing ENOMEM. Most processes don't deserve to starve the system completely.
Has this application already allocated 90% of the memory? Has it been steadily growing the allocation without releasing much back? Well then, why let the system run out completely? Why not stop it at say 10% or 5% left?
For a server, perhaps, but I'd say even for a server it would be better that say the core application returns a 500 internal error or whatever than forcing the system to start killing random processes.
I don't know how common this is on the Linux/POSIX side, but in DOS and Windows it is usual practice to allocate some amount of memory at startup and use that for error-handling code, so that things like showing error messages will not cause any more allocations.
It's also hard to get a good idea of how much memory a program is really using, which makes setting reasonable limits for things tricky.
The best you can do is to have a small swap space, and alert when it gets to 50% and fix whatever. But then you have the problem of filesystem pages evicting anonymous pages to swap which ruins the utility of the swap usage as a gauge and/or drives you to configure way more swap than is reasonable. Although, maybe a bigger swap space and alerts on swap i/o rate might be ok, other than it's awful painful to have swap space of even 0.5x ram if you have 1TB of ram.
The middle ground is your own config unfortunately - cgroups can limit available memory, but you'll have to set it up by hand.
For the general case though, because of the way programs have been written, it is easier to have overcommit on.
Many databases rely on the overcommit being possible.
You cannot make reliable systems thinking like that. The point of resilience is __not__ to rely on correctness of kernel's advertised behavior nor correctness of your assumptions about it.
One cannot build highly available systems and networks on unreliable, lying software. Basic guarantees are required. If one has 10,000 systems and one loses power on even a portion of such lying systems without basic guarantees, no amount of distribution will guarantee data consistency. I'm no stranger to building very large distributed clusters, but those builds start with an operating system and software which do not overcommit and which are paranoid about data integrity and correctness of operation. In fact, I'm specialized in designing such networks and systems, from hardware to storage all the way up to application software.