I first ported a toy app that I wrote (http://www.instantwordsearch.com) to EC2 - as an exercise. I personally learnt a lot in the process. Now I use EC2 for everything I do. It is a LOT of fun (imho).
Interesting links: http://jimmyg.org/2007/09/01/amazon-ec2-for-people-who-prefe... http://jimmyg.org/2007/09/01/custom-debian-ec2-amis-from-xen...
This is the kind of thing we were trying to avoid and ultimately (again) decided on our own servers in a data center.
If http://www.instantwordsearch.com/ look up is painfully slow but http://www.instantwordsearch.com/backstage.html (statically served 103KB text) is very fast, then you can conclude that the problem is with my app. Otherwise it might be safe to blame EC2.
The problem with EC2 and other super-managed hosting solutions (Joyent comes to mind) is inconsistency: I am sure they're moving VMs around, and occasionally they end up on overloaded servers.
I am designing a little bit on the paranoid side and making automatic incremental backups in S3. Recovery script will automatically restore using this too. I've simulated failures and I know it is safe. (this is something you need to do anyway - there are no guarantees in life!)
I still haven't automated the DNS change that needs to be done when you get a new instance. I will be handling this soon. There are some cool hacks already out there if you google for ec2 dns
My current approach is more exploratory coz I am learning as I go. But, to answer your question, if the instance were to fail right this moment, I'll get a notification and I'll run a recovery script that I've set up. I won't have this manual-intervention method for long though.
The S3 piece of the equation brings it back up in usefulness...but I don't know that I'd want to rely on it exclusively for my web applications just yet, unless high volume storage was a core part of my problem domain.
Instances can reboot, under operator control or due to a crash, and still retain their existing hard drive storage. Only less frequent 'instance termination' causes a loss of hard drive contents.
True, there are no guarantees that any instance won't be terminated by Amazon or other system failures at any time, so you're supposed to have your own backup and persistence strategy. But in practice, full terminations are rare, and Amazon has even given warnings when system upgrades mean large instance turnover is expected. (Though, there is no guarantee they will do so.)
So the falsehoods above are "[data is lost] if your virtual server goes down for any reason" and both concrete examples given, "taking it down to install security patches" and "system crash".
I looked back on my logs and found out that the particular degraded instance was running since June without any problems or shutdowns. Not bad for a VPS...
I'm very aware that Amazon provides no guarantees, but I have to say that my confidence on EC2 increased after this incident.
However... we have our web servers in a colo. Ping times to EC2 aren't fantastic, so we use a traditional CDN for static content and a colo for application servers in order to keep things snappy.
My requests time out an awful lot:
Read timeout. Please try again later. If this persists please visit the AWS developer forums to see if it's the result of a known issue.
That and doubts about permanent storage are downers.