I'm sure you have all heard of the Chaos Monkey that Netflix runs? Well, I didn't actually even need to code a chaos monkey... All you have to do is run everything on spot requests. Eventually you will lose servers at unpredictable times because someone outbids you[1].
The typical spot request pricing is at least 3/10th of the price of an on demand. For instance a cp1.medium (4 cores 4GB RAM) costs $0.044 per Hour for a spot request. Compare that to on demand and it is $0.186 per Hour.
I bid $1 per hour for my spot requests across two zones in the same region. I group my servers and use ELB (Elastic Load Balancers) to route requests...
Typically, a spot request might last for about a week before it gets killed because the capacity isn't there. That's then when instances in my other zone takes 100% of the load temporarily. At this point, since I've lost an entire zones worth of servers, I have my auto scaling group fire up on demand instances until I can get some more spot requests fulfilled. Creating a setup like this took about a week, but the savings are enormous.
----
How is the data stored on this setup?
RDS(MySQL) handles all data that can be stored in a database.
Ephemeral Storage is used to store things that don't need to be persistant (ie transactional logs).
Sessions are managed through Redis.. If the redis servers die, then session handling is handled via mysql temporarily. (It's a lot slower but the mysql server is RDS so it's always running)
Elastic Block Storage volumes are automatically mounted to a single instance which is then set up as a NFS server->client in order to allow other servers to read from a particular mount point (ie.. A user uploads an image, and it's stored on the NFS mount. A different server reads the file and starts generating different dimensions, uploads all of the files to Amazon S3, and then deletes the original file on the NFS device).
The worst part about losing servers is when the memcached server dies, because I could lose weeks worth of cache storage. When this happens, I have to boot up several micro instances that take my "cache warming" list and basically start repopulating memcached again.
The entire system is designed to be redundant... I can kill every server and then run the initialization script to start up the entire stack. (It's basically lots of little cloud-init scripts[2])