Source: I've got hundreds
Source: I've got hundreds
Instance and host reliability aren't the issue here. The issue is that a disgustingly high number of enterprises rely on vSphere High Availability and Fault-Tolerance features to restart VMs or keep them alive when an OS hangs or the host fails, instead of architecting for high availability at the app layer.
To be entirely fair though, vSphere HA and FT are incredibly friendly to the bottom line relative to rebuilding apps that simply weren't designed for HA.
I thought it was the usual complain about AWS instances dying and needing to be replaced on regular basis. That thing is a myth. An instances can run for years without any issues, if noone stops it manually.
Sometimes.
And sometimes the hardware it was running on dies or there was a maintenance event. Now you have a message like this waiting for you:
http://stackoverflow.com/questions/34259924/instance-retirem...
Fail to notice? Bye bye instance.
That said, you're right, it's not any more common than hardware failure. But that's common enough. Keep a backup and don't expect your stuff to always be there no matter what.
"High Availability and Fault-Tolerance" would hopefully involve more than restarting the VM.
I mean, you have like bajillion people that solved the above "challenge" with a three-line bash script. And not one of those people would call it "High Availability and Fault Tolerance". It's just a small shell script.
vSphere FT replicates a VM while it runs so that it doesn't even notice the failure of one of the underlying servers.
How often is your hardware fucked up enough that you need to move to another machine?
Honestly, if that happens often, there is something wrong with your hardware or your hardware provider or something.
On a 900-servers fleet on AWS, yeah, sometimes it would warn me that "the server needs to be retired" or whatever. Then I stop the server, then start it. Sure, inconvenient, but happens maybe... once a month? In fact, the frequency is decreasing so maybe once every two months?
> vSphere FT replicates a VM while it runs so that it doesn't even notice the failure of one of the underlying servers.
That would be great, right? You paid some $ to VMWare and now your servers never go down?
Excuse me but, that did not happen.
Makes you think.
Thanks.
VMware HA runs quite a bit IME for a variety of reasons (network failure, storage issues, etc). "The network is reliable" is a classic fallacy of distributed computing.
More often used is Vmotion and DRS which transparently moves VMs around physical hosts at runtime with no downtime and very minimal performance hit . Servers, switches, and fabrics need maintenance, firmware upgrades, hypervisor OS patches, etc. in most data centers.
"That would be great, right? You paid some $ to VMWare and now your servers never go down? Excuse me but, that did not happen."
VMware FT is used in almost every major company on the planet that runs virtualized database instances that can't go down. It is serious technology.
AWS has mopped up the IT industry to date, and this is them going hard after Microsoft which is making inroads.
That said in recent years it's much less of a problem , though the fixes tend to require a lot more expenditure for provisioned IOPS, dedicated hardware instances and 10 gigE network instances. But then nearly any legacy shit app can run in AWS ...unless it needs some crazy specific configs like live hardware assisted disk replication, or VMware FT levels of resilience.