How we spent Friday night coming back online before Instagram and others
blog.fitocracy.com
blog.fitocracy.com
I had a site that was affected by another one of Amazon's outages a while back, and here was my disaster recovery plan in its entirety:
a. go to sleep.
There's a reason you farm things like this out to Amazon in the first place. They have a big team of smart people whose only job in life is to keep your stuff alive, or scramble like mad to bring it back up if it goes down.So long as your site knows how to start up automatically when the box turns on, there's really not a lot you need to do in a situation like this.
If forty percent of the internet is down, and you're part of it, your users will probably understand. They'll expect you to come back up when the rest of the internet does. If you do manage to come up a bit earlier, you might get a shrug and a "cool", but it's probably not enough of a win to cancel Christmas.
The returned record, of course, has a TTL, and the ALIAS mechanism will not alter the TTL of the aliased data, so in the case of an ELB you are talking a TTL of one hour. There are no magic bullets to the distributed cache expiration problem.
(The other comments in this article about TTL seem quite confused, though, so this explanation might not actually have helped. Even if you have a 20-year TTL, you are going to see changes immediately from clients that do not have the data cached anywhere on their path to the origin.)
(In particular, there is no difference at all to switching the ALIAS record out for an A record with regards to your TTL: if the user has the target of the old mapping cached they will use it, otherwise they will get the new one. It isn't really "getting away with" that behavior, and it isn't sue to Amazon's DNS being special.)
(they do this in an effort to reduce load; whether it is truly effective on today's hardware is arguable)
Even before you can afford a part-time DevOps engineer I highly recommend automating your system administration as much as possible. When you say you should script your installation of an application server I recommend doing that with Chef!
This will allow you to quickly bring instances online on about any cloud. HP, AWS, Linode, or your own self-brewed OpenStack cloud. You will still have challenges with your persistent data but it's a lot easier to breathe and act quickly with your persistent data knowing you'll have content servers to access that resource.
edit: I have no affiliation with Opscode - we aren't even a paying customer we use their free server. I am sure Puppet or any system administration tool will get you similar milage.
OVH have an android app so I can scale to a new server from the push of a button on my phone. :D
Plus they cost peanuts compared to amazon.
I has fall over in other data centres too but OVH makes me the happiest.
And I think you hit the nail on the head - 'regions' are for regional fault tolerance, while AZs are for within-a-region fault tolerance.
We waited for stuff to come back up. Our relatively simple application required far fewer servers than more complex services with millions of users.
We're awesome!
Do you think there's some middle ground of using AMIs but also using puppet somehow, so you make new AMIs as a perf optimization but keep puppet config up to date? TBH it's something I've only casually wondered about. But maybe it's what we both need. Having a puppet config would mean you can launch on basically any provider.
However, we're probably going to make a deployment script (we use fabric) that builds the AMI from scratch. Then, when we need to update the AMI, we just update the fabric script and run it to make the new AMI. That way if we ever need to make AMI's in a different region, we can just run the fabric script in that region.
This gives the best of both worlds - version controlled configuration files, and an automated process for making new AMIs. Instances also boot quickly, as nearly all of the configuration is baked-in to the image.
Using EC2 tags gives you even more room for automation here.
It also sounds like you could pretty easily make your AMI creation step a job in your CI software.
1) Set up a test environment using puppet or whatever that boots up an image instance and deploys your whole stack to your cloud provider. 2) Deploy your release candidate to this environment. 3) Run some smoke tests, acceptance tests and maybe even a few performance tests against this environment. 4) If all the tests pass store your image somewhere (S3, GitHub) also make sure to tag the image with your release candidate version. 5) You can now deploy this change-set to your production environment.
When things go wrong you will have fully configured servers available almost immediately.
Now granted this is potentially a very long running process but you could simply set this up as a nightly build.
It relies on several cloud computing providers, to prevent such drops to happen.
The pricing is even lower Amazon or any other computing provider! (we buy larger clusters which costs us less)
You should THINK about your TTL values long in advance of a problem. You should also THINK about having a backup instance running (or at least ready to boot, if an hour of downtime is perfectly OK for your users).
More discussion here: http://news.ycombinator.com/item?id=4181918
Sounds to me like they learned a lot, and are sharing what they've learned to help others. What's really to gain from shaming them over it?
I would be a lot more impressed if they were asking questions about how DNS really works and what they should do to avoid problems in the future. That would be cool.
The reason you got your site online before Instagram and others is because they have a lot of infrastructure and moving pieces as a result of being extremely popular. Obviously Fitocracy doesn't share those characteristics.
That said it is unacceptable for ANY site to go down simply because you lose power in a data center.