At scale, server failures are going to happen. Maybe it's a power outage. Maybe it's a failed drive. Maybe it's a kernel crash. Only one of those can be remedied (not even prevented) with careful debugger investigation, so you're going to have to do the work to be able to handle such failures gracefully. At that point, unless the leaks are causing outages so rapidly that the machines can't handle, then rebooting becomes not that big of a deal, while wasting senior eng days trying to track them down is a real cost.