The problem here was that something that is known to fail for all sorts of reasons (network IO) was happening without such a timeout. Or with a timeout with a failure mode that it never happens (yikes). That's a design problem and even something with a very small chance of happening is extremely likely to actually happen at some point with a product that is this widely used.
This stuff is hard of course and I end up addressing issues related to his once in a while. The fix is usually to surround such code with defensive measures such as timeouts, retry mechanisms, telemetry, logging, etc.
The additional question/learning is why they never noticed this happening before. Because it probably did; they just never noticed because the very thing that would have told them was actually hanging. People killing an application for whatever reason is something that you'd want to know however.