AWS outage brings Netflix down for some devices on Christmas Eve
engadget.com
engadget.com
Netflix, to me, is a big collection of links to stuff in CDNs (which, got a first approximation, never go down; Akamai is essentially unfadeable, highly resistant to all forms of outage because it's a trivial replication problem.)
I rarely if ever care about the recommendations engine, new content in the submission queue, etc.
Yes, there's an authentication problem, but this is also trivial to replicate, and it's fine if "is a valid subscriber" goes even a month out of date.
Essentially, even if AWS goes down completely, the Netflix client should be able to get a static list of movies and show them. Maybe that's my queue, maybe that's the top 10 per genre, whatever.
In times of degraded operation, show me something.
(It wouldn't have helped in this case, but its a general annoyance I have.)
Basically after nearly 15 hours of downtime I'd consider that beyond unacceptable. With physical hardware even on Christmas Day you could have replacements well before that and spun up... just saying.
When I worked at a hosting company, we'd have at least 5 shared hosting servers (with ~75-200 websites on each) go down each day, and yet we had a reputation for good uptime, because comparatively, we DID have good uptime.
I think the problem is that most people think about uptime wrong. Uptime is a compromise; a trade-off. You can pay for more uptime with diminishing returns after about 98%. Running one dedicated server colo'd with decent fault recovery systems in a decent datacenter will probably get you 100% uptime most of the time, until it doesn't, thus ~98%.
If you're going to attempt to get beyond power failures, you'll need a second server (or "instance") somewhere. If you need the capacity anyway, this might not double your costs, but you may still have unused headroom. You could buy flexible computing power by using shared hosting ("the cloud") or whatever, but it's the same problem.
Once you get into a state where you have a global business, customer demands, supplier issues, vendor lock-in, etc, it becomes a numbers game. You can hire (more|better|famous) devs and possibly get more uptime. You can test more (and slow down feature releases) to get better uptime. You can break stuff (and pay for the recovery + lost face + downtime) to decrease downtime. Everything is a trade-off, and right now it makes sense to chase about "four nines" of uptime -- 99.99%.
Four nines is 4.32 minutes per month -- four minutes and nineteen seconds AT MOST once a month. This is very possible for many large services and while it does have cost overhead, it's manageable.
TL;DR don't go chasing waterfalls (100% uptime.) It's not practical. It may be possible, but it depends on how long. 100% uptime for 10 years would be a pretty damn lofty goal. In my opinion, it's much more important to recover quickly and gracefully with awesome communication with your customers than it is to be up 100% of the time. 100% uptime goes unnoticed, but consistently great customer service does not.
I guess redbox is still working.
I wonder what percentage of the population will have never owned a DVD player in the next generation.
If netflix owned their own hardware and could reach out and touch it, would this have happened?
Probably the same amount of people that have never owned a record player or, for the younger people, a CD player.
> If netflix owned their own hardware and could reach out and touch it, would this have happened?
Yes, outages happen whether you have the hardware in your own hands or have it hosted by a cloud provider.
I'm not questioning if there's a good reason that it's done this way, but that reason just not obvious to me. I would have designed such a system where there is one endpoint which all clients hit, regardless of platform. But I have never designed the world's largest video streaming infrastructure.
They then put all developers on call, and force them to write code that can recover from faults by trying different methods to break it.
tl;dr netflix systems are chaotic, because chaotic systems tend to tolerate failures.
Since many of these other platforms also use proprietary formats they end up having to maintain a range of different proprietary streaming servers, many of which presumably need to interface with their respective mothership platform services (XBox live, PSN, Roku).
Oh, and recently they have begun building their own CDN[1] which may add further to the diversity.
Different for each device based on how much power the device has. TVs have nothing for CPU power in most cases.
- SSL for "free"
- If you use Route 53, it's a single DNS lookup (no CNAME hop)
- If you're using a VPC you can use an ELB to face the public Internet.
- (Slightly simplier) auto-scaling.
- Redundancy across AZ's
Many/All of these things you could achieve yourself, but if you were using EC2 resources you'd probably find it reasonably expensive.
http://blog.rightscale.com/2010/04/01/benchmarking-load-bala...
I say these because I don't want people to get a sense of false security when deploying to multi-AZ behind ELB.