I think a big takeaway from it is that designing systems which are failure-free is a fool's errant - no matter how hard you try, you can never get rid of 100% of the bugs.
Instead, make it failure-tolerant: sooner or later every part of the system will break, so it should be constructed in such a way that it can gracefully recover from failures, and even operate with some parts of it unavailable. Crashes are expected, so the system is designed to handle them properly.
A fragile thing is like a wine glass. Once it breaks, it cannot be restored to its original state and especially not made better than before.
However, if you're talking about patching the system while it's still running, check out "Stop Writing Dead Programs" from Strangeloop '22: https://youtu.be/8Ab3ArE8W3s
You don't get what you don't pay for™
End-to-end and integration tests are much more helpful. But even then, they won't look at operational concerns like backup and recovery.
For software, that would mean practices like Chaos Monkey. If your production system stays up while an external process is constantly killing processes and deliberately corrupting memory and files then you have good confidence of riding through unexpected failures.