Lessons Netflix Learned from the AWS Storm
techblog.netflix.com
techblog.netflix.com
http://techblog.netflix.com/2011/04/lessons-netflix-learned-...
Though their post today makes it sound like they could have (and maybe did) fail over to an entirely different region, but their mechanism for doing so isn't automatic and took longer than expected.
There are ways to do it that take seconds (we do that at Quantcast using anycast), there are ways that take minutes (using DNS failover, which is readily available), and there are ways that basically take forever (Netflix guys - feel free to contact me).
It's pretty clear they were aware of the problem, and if they had even the simplest and most basic DNS control, they could have moved people over minutes after realizing the problem.
Being able to move /millions of requests per second/ (this was a Friday night!) on /thousands of servers/ [2] to a different /thousands of servers/ involves more than just routing requests to the right IP. Like having a separate set of /thousands of servers/ capable of handling that much failover, for one. Plus, if you had read the link, what they just started talking about a year after moving to AWS was moving across regions; they're already set up to fail entire availability zones over to other availability zones in the same region.
Neither architecture at scale nor employee experience magically appears from nowhere. It's not even been two years since they moved from a data center to running the site on AWS. Considering how much they've had to learn, and the kinds of tools they've had to build for themselves to manage their kind of scale (read their tech blog!), I think the cheap insults are unwarranted.
1: http://news.ycombinator.com/item?id=4208134
2: http://techblog.netflix.com/2012/02/fault-tolerance-in-high-...
Even a child knows that when the power fails, you can't turn on the TV. This isn't specialized technical knowledge.
EDIT: I'm not going to keep responding to your comment below. I'm certain that if I were involved in the design of Netflix' infrastructure, they would be able to survive problems that affect whole regions. (AND I DONT SEE THAT NETFLIX NEEDS 6 SECOND FAILOVER, THERE ARE MUCH EASIER WAYS TO DO IT IN MINUTES).
EDIT: My repeated, emphatic comments are intended to serve a purpose. Everyone should be aware that this is a real problem and you need to plan for it, and it's pretty clear from the discussions on HN that people are surprised by the Amazon downtime. I personally think Amazon does a fantastic job as it is, and Amazon's reliability issues are not an excuse for the downtime of their customers.
You're basically saying "if I were in charge, Netflix would've been able to fail over to another region in 6 seconds". But you don't even have a fraction of the background info required to say such a thing. It doesn't matter that you've done reliability at Quantcast; Quantcast is not Netflix.
It sounds like they came to a different decision than Quantcast about the importance of being able to do so.
Which is more likely - that the thought of an AWS outage like this never occurred to them, or that they judged the specific remedy needed to overcome it not a high enough of a priority to have it available within the first years of their AWS usage?
AWS doesn't support multicast or anycast, and the AWS EC2 control planes were so hosed it was impossible to recover in any meaningful way. Certainly, the fault was both AWS and Netflix, but both are learning from their mistakes.
Back to AWS, the control plane failures are concerning, killing the ability to deploy new resources. It's expensive to have this kind of capacity pre-deployed and I'm sure it bit many folks.
I'm happy to help anyone who is serious about uptime, just email me. There's a reason that I post under my actual name, and there's a reason that I make my contact information avilable.
Chill out.
In this case we had some bugs, we should have had a two minute increase in error rate as a third of the clients retried, then the dead instances would have been out of traffic. That's what happened in the previous power outage, where fewer instances went down, and it didn't trigger this bug.
Also, as someone who was a user of your globally available service, I can tell you that while it might have been UP all the time, it certainly had no problems losing data all the time too. Some months there was just simply no data for reddit at all, even though we were sending the service more than a billion data points.
So we could sit here and sling insults all day, or you can operate under the assumption I do -- that each of is a competent engineer who works for a business that has to make decisions and tradeoffs between costs and reliability.
My understanding: Streaming is offloaded to a CDN. The content & user databases change very slowly and thus are very amenable to caching. 'Current position' syncing is very non-critical and so can be done with e.g. writes to independent Redis instances (or even memcache - it really isn't very important!) Ratings, recommendations and the queue seem like the tricky ones, and while I don't think the throughput is particularly high, because they are all per-user this is "trivially shardable" if you do outgrow a single SQL database.
The big question for me is understanding why streaming breaks, because that should all be served from systems that use read-only data? (where read-only = cacheable for at least a day without significant negative consequences)
I think a better understanding of Netflix's technical challenges would serve everyone well here!
http://techblog.netflix.com/2012/04/netflix-recommendations-...
In parts one and two, it's explained that just about every list generated for users is an up-to-date set of recommendations. You point out yourself that recommendations is one area where caching it's truly viable (at least, not in a traditional sense).
I'm definitely not the right person to go into all the details (nor do I think such a discussion would be prudent on HackerNews)--but I wanted to weigh in quickly that there's a lot of stuff served that goes way beyond the notion of a "static" content that's trivially cached.
I feel like the quality of the Netflix recommendations is not stellar, and if that's because you're constraining yourself to what can be calculated in real-time, I'd willingly trade-off having "perfect" real-time recommendations in favor of better recommendations tomorrow (with the full model). Even if you do try to update recommendations in real time, aren't they easily cacheable if you can't keep up? (Well, as easily cacheable as any dataset on 25 million subscribers can be...)
Though I'd love to see the monitoring solution open-sourced :-)
Anyone who reads HN can see that a minimum uptime strategy for Amazon is to failover across regions. Each time there is a major AWS outage, we hear about HN readers whose service was affected even though they spanned availability zones within a single region. But to date, Amazon's regions have operated independently.
That observation is not dependent on knowledge of Quantcast (which is incidentally, far more than a write-only system), or the other production systems I've built in the last 35 years.
(I'll follow up by email about your support questions)
(I'll assume it will be a routing problem or a software problem.)
"Don't panic. You are using a backup datacenter. Some very recent queue or account changes may be missing, and some changes you make tonight may be lost. We are working nonstop to resolve this and appreciate your patience"
When stuck, just change the requirements.
If the netflix business fails, they would have giant valuable datacenters leftover. Instead by relying on the cloud, they are "all-in" on serving movies. The movie and tv studios have giant leverage here, they can easily make or back a competing service and users will go where the content is. Is their strategy for being on the cloud really, "it's easier than doing it ourselves?".
It's not like Amazon where they're providing infrastructure to other companies. I remember someone else on HN pointing out that there aren't many non-adult video providers the size of Netflix/YouTube/etc that aren't already rolling their own solutions or served by companies like Brightcove.
You should keep your logics dumb.
It sounds like there was some technical debt to that implementation, but hey, I for one am glad they gave us some insight into what happened.