Netflix: Lessons We’ve Learned Using AWS
techblog.netflix.com
techblog.netflix.com
As Netflix continues to scale, these changes will make managing that growth much easier.
A lot of you seem to take this post as being negative against AWS architecture. I take it more as a good collection of common things that you need to watch out for in distributed environments, specifically the dangers of assumptions within your current infrastructure which may change dramatically as you scale.
The best way to test the uncommon case is to make it more common.
Some great ideas to be gleamed from the paper you've provided - thanks!
Chaos Monkey is just a script that runs "ec2-terminate-instance" commands on node in certain security groups at a certain rate, and boots machines at the same rate.
Of course, the devil is in the details, but at a high level this wouldn't be too specific to the domain the infrastructure is built around.
Agree that implementing similar system for a single AWS account is not going to be very difficult.
Actual URL is http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.90.... and that is a direct link to a PDF.
I should have noted that my link was a PDF, though. My apologies for that oversight.
[1] http://www.google.com/search?q=%22[1]%20http%22+site:ycombin...
While the likelihood of failure (or added latency, impacting upstream changes, etc.) is greater in large-scale distributed environments for which you do not control vs. your home-grown datacenter, those scenarios are just facts of life in distributed environments.
An awesome side effect of hosting an app in a cloud environment is that you must face up to those fallacies immediately or they'll eat you alive.
If a company has a short runway (ie. not a lot of cash in the bank) however, hosting in the cloud means essentially renting capacity for which you do not need to maintain. Sounds good to me.
It's a fun pendulum to watch swing though, I'll admit. Some companies, with small runways, host in the cloud because it's rented capacity without too much administration. They may grow wildly, when they notice they are paying the "cloud tax" and could considerably save money by hosting themselves. They grow some more, see how much money and attention they need to put into infrastructure and how slow it is for their business to increase capacity. In order to keep up with growth and maintain focus on their core business, they move to the cloud.
Netflix is a going concern, and has been for quite some time. Even the streaming business is a few years old.
It may be an up front investment since the full benefits may take some time to accrue, but it is an investment made based on a great deal more information than a typical startup has.
Further, it sounds like an investment that is providing them with an immediate and valuable benefit, they now have a foundation for ramping up their streaming business. This is going to be a big growth area, and it's is one that they are well positioned to succeed in. Or, you could flip it around, failure to secure a strong position in the streaming business will be the death of the company, and a squandering of the business equity they have been building, since their founding, to be well positioned for this transition.
And finally, if you have doubts about the value of making up-front engineering investment in order to reduce forecastable operating costs (and avoid what is probably an even larger up-front investment in datacenter buildout), well, what are you doing here?
Actually, that was my first reaction, but after thinking for a moment, that isn't really a reliable way to test. If you make changes to something, you don't know for sure if the chaos monkey hit while you were testing a certain thing or not. Proper unit tests would seem to be a lot more useful.
OK, it ain't perfect. Things usually doesn't fail in a uniformaly distributed way and you can't be sure you'll see a different kind of fail on production, but it sounds useful(yeah, and pretty cool) nonetheless.
"SQLite responds gracefully to memory allocation failures and disk I/O errors. Transactions are ACID even if interrupted by system crashes or power failures. All of this is verified by the automated tests using special test harnesses which simulate system failures. "
Chaos Monkey is meant to force apps to go into failure modes "in the real world". If they see anything unexpected happening after a service is killed or degraded, they can investigate.
It sure beats having a service go down unexpectedly for the first time six months ago and not have tested scenario in production.
This excellent paper[1] by James Hamilton (then MS, now AWS) recommends never doing clean shutdowns on applications - just kill them. Unless you have a lot of persistent state to manage this is a great idea. If it takes hours to migrate multiple TB off your storage host, maybe not such a good idea.
1. http://www.usenix.org/event/lisa07/tech/full_papers/hamilton...
The tone of this post indicates to me that the criticism and problems experienced by Netflix with AWS are understated, which I can understand given their position as a flagship AWS customer, etc.
Another way to say "If it ain't tested, it's broken".
Failures were always going to happen, even in their own datacentre. What they have now is a more fault-tolerant system which should have less downtime overall.
Otherwise, good idea. It forces you to think about the perils of distributed environment from the very beginning, as opposed to leaving it to be an afterthought.
No need to simulate traffic for testing purposes. Here's our actual traffic. All of it.
Nice.
I say "feel" advisedly. I don't have inner knowledge. It's just that the explanation up to this point doesn't seem all that compelling; for what seems like rather dubious benefits it seems like they've taken on an awful lot of risks they have little ability to manage themselves.
http://techblog.netflix.com/2010/12/four-reasons-we-choose-a...
The problems [Amazon] are trying to solve are incredibly difficult ones, but they aren’t specific to our business. Every successful internet company has to figure out great storage solutions, hardware failover, networking infrastructure, etc.
Their third argument for virtualized infrastructure was "We're not very good at predicting customer growth or device engagement", which strikes me as a compelling problem to solve. Having to change their approach to managing service dependencies and failures doesn't seem like too high a price to pay.
This is not Netflix's core competency either. This is the core competency of a CDN. If you have problems like this, you should almost certainly hire one, and architect accordingly.
Netflix has long since done so:
http://blog.streamingmedia.com/the_business_of_online_vi/201...
netflix actually uses many cdns, including akamai and l3.
Now with AWS, they can spin up additional instances at will.
It's just a matter of alternatives. There's frankly no other IaaS out there that can match the features of EC2 (eg, EBS, elastic IP). These guys are unstoppable, they come out with more features every week.
a) Ability to map an elastic IP to an ELB. The CNAME thing it uses now has way too many drawbacks [1].
b) Ability to make an RDS instance a slave to or a master for a normal MySQL instance. This would make it possible to use RDS as a backup to our ordinary DB infrastructure using normal MySQL replication (and eventually, vice versa)
c) Retention periods for EBS snapshots. You can get this yourself by writing a few simple scripts, but it would be really nice if snapshots could simply be labelled "delete after 90 days".
d) A cross-availability zone, synchronously replicated EBS volume. This is probably fairly specific to our use case, but it would be neat if AWS natively provided something like DRBD.
[1] http://blog.pagerduty.com/2010/08/31/load-balancers-need-sta...
Let us add dev pay instances to an ELB.
More ram.
Elastic private ip addresses.
Change security groups of running instances.
I have a lot more, but those are my big ones.
I did have this one question, being a guy with an IT background: they expected stability? Really? I always expect host/app/system failure, and am pleasantly surprised when it doesn't happen.
With your own IT infrastructure, you'd logically put the focus into improving the reliability on your hardware stack, rather than just focusing on handling those issues in software.
Netflix is a sizable and successful business. It's more than a decade old. Their streaming business launched only 4 years ago, more than 6 months before EC2's general availability, and almost 2 years before EBS, and it would be even longer before any of it was a reasonable choice for a company of their size, launching a new service building off their existing customer base, that had to succeed to allow the company to navigate a major transition that the company had been planning for for years.
The similarities and differences in their experiences would be interesting and probably informative, particularly if considered against the vastly different starting points.