"but we have lost machines, connectivity, seen kernel panics, EBS failures, SSD failures, etc., your usual day in AWS " <=== This I wish more people realized that is a day to day reality if you are in AWS at scale.
https://cdn.chrisshort.net/How-Complex-Systems-Fail.pdf
Basically once a system is complex enough some part if it is always broken. The software must be designed from the assumption that the system is never running flawlessly.
Or are you saying that AWS is particularly unreliable at scale?