Hardware Failure on EC2
forums.aws.amazon.com
forums.aws.amazon.com
Hardware is fairly reliable nowadays (whether you're hosting your own or using the cloud provider's) but failure still happens.
The bonus that you get by using a cloud provider is that you get API machinery and tooling to automatically handling most failures.
The malus is that depending on your level of paranoia (eg: architecting for multi-region) redundancy will get very expensive very soon.
But in the end... meh.
Despite what "evangelists" or "detractors" will say, there's no free lunch.
GCP has hardware failures, and if there's a bad enough one it impacts your instance, but you can have instances that live forever and they just move it from physical machine to physical machine and it mostly works.
I was very skeptical about this, having spent 5-6 years on AWS and got used to the idea that AWS might just blow your instance away (normally with notice, sometimes not though if the hardware hosting it had some catastrophic failure). GCP just says, "Don't worry about it, things will get moved" and the servers in question have long lived connections, they aren't just simple stateless HTTP API servers, and it works 99.9% of the time (that last 0.1% of the time is a real pain to debug).
We still treat our machines as cattle instead of pets, but not having to constantly deal with cattle roaming off the reservation is nice.
Yes your code should handle an instance dying. But if 98% of the time you can live migrate and not... hey why not.
AWS has had a partnership with VMware for a while now, so I would imagine/hope that they could still do live migration, even when running in the AWS cloud.
That does make me wonder why AWS doesn’t implement the same functionality native, though.....
The result is, as one engineer on our team says, you basically run chaos monkey but amazon pays you to do so.
1: https://www.vmware.com/products/vsphere/fault-tolerance.html