Amazon EC2 currently down. Affecting Heroku, Reddit, Others
status.aws.amazon.com
status.aws.amazon.com
The happy medium is probably splitting database (master/slave at least) and cdn (if needed) and some other services (AAA? logging?) out, and then having 2+ front end servers with load balancing for availability.
In the first 99 comments of this page, average comment text size is 231 bytes. Counting all comments in articles on the front page right now, there's 1678 of them, making somewhere around 388kb of comments for the past 12 hours.
So for safety's sake round that to 1mb/day and multiply by site age (5 years).
That gets us 1825mb, projecting forward it's difficult to imagine a time when a single recent SSD on a machine with even average RAM wouldn't be able to handle all of HN's traffic needs. Considering the recent beefy Micron P320h and its 785kIOP/sec, that could serve the entire comment history of Hacker News to the present day once every 2 seconds, assuming it wasn't already occupying a teensy <2gb of RAM.
Even if Arc became a burden, a decent NAS box, gigabit Ethernet, and a few front end servers would probably take the site well into the future. Assuming exponential growth, Hacker News comments would max out a 512GB SSD sometime around 2020, or 2021 assuming gzip bought a final doubling.
You wouldn't notice if it weren't for the use of closures for every form and all pagination, every time the process dies all of them are invalid (except in the rare case that they lead to a random new place!).
There's no database, everything is in-memory loaded on-demand from flat files. That wouldn't be so bad except that it's all then addressed by the memory locations rather than the content identifiers! There can be only one server per app, and to keep it real interesting PG hosts all the apps on the same box, during YC application periods he regularly limits HN to keep the other apps more available.
Nothing has changed in the stack. Robert has discovered and eliminated a series of bottlenecks, causing performance to oscillate about tolerable. Finding bottlenecks is not trivial, because Arc has zilch in the way of profiling, but fortunately Robert is good at this sort of thing.
(this isn't to downplay the challenges faced by scaling a site with the amount of traffic HN gets)
A magic bullet it isn't.
Slays millions with a single round.
http://media.amazonwebservices.com/AWS_Cloud_Best_Practices....
“Be a pessimist when designing architectures in the cloud”
http://media.amazonwebservices.com/AWS_Web_Hosting_Best_Prac... “As the AWS web hosting architecture diagram in this paper shows, we recommend that you deploy EC2 hosts across multiple Availability Zones to make your web application more fault-tolerant.”
http://media.amazonwebservices.com/AWS_Operational_Checklist...
“We have deployed critical components of our applications across multiple availability zones, are appropriately replicating data between zones, and have tested how failure within these components affects application availability”
So it seems that the only real benefit to utilising cloud services is to make scaling up easier and save money.
"Building your own" is also something very few people (including Amazon itself up until fairly late, probably after the IPO) do: you can use a managed hosting provider (very common, usually cheaper than EC2) or lease colo space (which doesn't imply maintaining on-site personnel in the leased space: most colos provide "remote hands"). You can still use EC2 for async processing and offline computation, S3 for blob storage, etc... or even S3 for "points of presence" on different US coasts, Asia/Pacific, Europe, but run databases, et al in a leased colo or a managed hosting provider.
Yes, these options are more expensive than running a few instances in a single EC2 AZ: but that's the price of offering high availability SLA to your customers. It's a business decision.
E.g. I'm currently about to install a new 2U chassis in one of our racks. It holds 4 independent servers each with with dual 6 core 2.6GHz Intel CPUs, 32GB RAM and a SSD RAID subsystem that easily gives a 500MB/sec throughput.
Total leasing cost + cost of a half rack in that data centre + 100Mbps of bandwidth is ~ $2500/month. Oh, and that leaves us with 20U of space for other servers, so every additional one adds $1500/month for the next 7-8 or so of them (when counting some space for switches and PDU's). Amortized cost of putting 2U with 100Mbps in that data centra is more like $1700/month.
Amazon doesn't have anything remotely comparable in terms of performance. To be charitable to EC2, at the low end we'd be looking at 4 x High Mem Quadruple Extra Large instances + 4 x EBS volumes + bandwidth and end up in the $6k region (adding the extra memory to our servers would cost us an extra $100-$200/month in leasing cost, but we don't need it), but the EBS IO capacity is simply nowhere near what we see from a local high end RAID setup with high end SSD's, and disk IO is usually our limiting factor. More likely we'd be looking at $8k-$10k to get anything comparable through a higher number of smaller instances).
I get that developers like the apparent simplicity of deploying to AWS. But I don't get companies that stick with it for their base load once they grow enough that the cost overhead could easily fund a substantial ops team... Handling spikes or bulk jobs that are needed now and again, sure. As it is, our operations cost in man hours spent, for 20+ chassis across two colo's is ~$120k/year. $10k/month or $500/per chassis. So consider our fully loaded cost per box at ~$2200k/month for quad-server chassis of the level mentioned above with reasonably full racks. Lets say $2500 again to be charitable to EC2...
This is with operational support far beyond what Amazon provides, as it includes time from me and other members of staff that knows the specifics of our applications, handles backups, handles configuration and deployment etc.
I've so far not worked on anything where I could justify the cost of EC2 for production use for base load, and I don't think that'll change anytime soon...
If you are very database heavy, and you want to be able to replicate that to the cloud in real time it does get expensive, but if you can tolerate a little downtime while the database gets synced up and the instances spin up that's cheap too.
They recommend that people failover to other availability zones but no one puts any effort into doing it then they get annoyed when a datacenter goes offline.
Its not Amazons fault that you didn't make your service failure tolerant - its your fault!
It reminds me of my father: I used to interrupt him with "I know dad!" when he was chastising me. His response was simple "If you knew, then why did[n't] you do it?"
We had servers in the bad zone and started having load issues. When I went to use the cool cloud features that are made for this, the entire thing completely fell on its face. I couldn't launch new EC2 servers either because the API was so bogged down, or the new zone I was launching in was restricted because of load.
Basically, the thing that nobody keeps in mind when they think it's so cool that you can spin up servers to work around outages is that EVERYONE IS DOING THAT. This is Amazon's entire selling point and when it comes to doing it, it doesn't work!
We were lucky to get some new servers launched before the API pretty much completely went down. They started giving everyone errors saying request limit exceeded. The forums were full of people asking about it.
ELB, Elastic IP, and other services not associated with a single availability zone completely failed. I keep seeing comments saying that if people designed their stuff right, they wouldn't have an issue. That's just completely bull, AWS has serious design flaws and they show up at every outage. It's NOT just people relying on a single zone.
The main systematic issue in EC2 is EBS, take that away and it will almost completely remove downtimes.
Where getting a bit from disk to memory used to be: platter -> diskcontroller -> cpu -> memory,
now with SANs & NFS & virtualized block storage, it's: platter -> diskcontroller -> cpu -> memory -> nic -> wire -> switch(es)/router(s)/network configs(human config item) -> wire -> nic -> cpu -> memory.
Not to say that centralized storage doesn't have its benefits, but now the scope of isolation has drastically increased, which when considering the combinatorial possibilities of failure in the prior scenario vs the latter, the latter has a significantly larger chance and mode of failure that is significantly more difficult to programmatically automate failover.
TLDR: With amazon, the scope is isolation is the datacenter. To be on amazon, one must architect and design at the scope of handling failure at the datacenter level, rather than at the host or cluster level.
I actually asked the AWS Premium support regarding the ELB multi-AZ issues, in order to actually make things easier for everyone. This is the answer I got:
"As it stands right now, you would need to make a call to ELB to disable the failed AZ. It may be possible for you to programatically/script this process in the case of an event.
Going forward, this is something that we would like to address but I don't have any ETA for when something like this might be implemented."
For the places that truly care about reliability and have the technical staff to make informed decisions, they should understand the limits of reliability with various architectures. As I mentioned before, one of the tenant of reliability is isolation. When the scope of isolation is increased (e.g. single host vs multi host), one must also handle failures at that scope. Amazon isolates at the datacenter level. So should those utilizing Amazon's offerings.
Take care when treating a high-availability set up such as this as a back up: if you are replicating all the changes between 2 database servers and an application error (e.g. not a database outage) causes some kind of data corruption, you are hosed if the corruption replicates and you don't also have some previous "snapshot" of the data that you can roll back to.
N. Virginia is both cheaper and closer to the center of mass of the developed world. I'm surprised Amazon hasn't managed to make it more reliable.
[1] http://huanliu.wordpress.com/2012/03/13/amazon-data-center-s...
This means that if you were trying to make changes to your EC2 instances in the West using the GUI, you couldn't, even though the instances themselves were unaffected.
I get tired of the snipes from people that "well, you're doing it wrong", as if this is trivial stuff. But if Amazon themselves aren't even making their AWS console redundant between locations, how easy/straightforward is it for anyone else?
To what extent is this just "the cobbler's kids have no shoes?"
We're not talking about a leap in order of magnitude of complexity here—just simple management of common human behavioral tendencies in order to promote more reliability. "The problem is inherently complex" is always true and will always be true, but it's no excuse for not designing a system to gracefully handle that complexity.
You're close. Put another way, "inherent complexity is the problem."
What I mean by that is, the more your system is coupled, the more it is brittle.
Frankly, this is AWS's issue. It is too coupled: RDS relies on EBS, the console relies on both, etc. Any connection between two systems is a POF and must be architected to let those systems operate w/o that connection. This is why SMTP works the way it does. Real time service delivery isn't the problem, but counting on it is.
Uncouple all the things!
us-east-1 was 11 different datacenters last time I bothered to check.
us-west-2 by comparison is two datacenters. The reason west-1 and west-2 exist is because they are geographically diverse enough to prevent low latency inner-connections (and also have dramatically different power costs so they bill differently).
Edited: [1] http://www.datacenterknowledge.com/archives/2012/06/30/amazo...
Weigh this against the estimated costs of your application going down occasionally. It's really only economical for the largest applications (Netflix, etc.) to build these systems.
Every day I have to build basic redundancy into my applications I wish that we could just go with a service provider (like Rackspace / Contegix) that offered more redundancy at the hardware level.
I know the cloud is awesome and all, but having to assume your disks will disappear, fail, go slow at random uncontrollable times is expensive to design around.
If you don't have an elastic load, then the cloud elasticity is pointless - and is ultimately an anchor around your infrastructure.
two c1.medium, which are very nice for webservers, are enough to host >1M pageviews a month (wordpress, not much caching) and cost around $120/mo each, effective $97/mo if you prepay for 12months at a time via reserved instances.
Also, What is happen in cloud is stay in cloud because nobody can able reproduce outside of cloud.
(And many other relevant quotes.)
What do you do when Amazon is down other than sweat?
2. Calculate the cost of moving to the Oregon AWS datacenter.
3. Reassure your investors that outsourcing non-core competencies is still the way to go.
4. Try to, er, control, your inner control-freak.
;)
But better and faster than Amazon?
I'd rather spend three hours at home saying "Shit. Well, we'll just wait for Amazon to fix that", than dropping my dinner, driving to the datacenter, and spend three hours setting up a new instance and restoring from backup.
"Source of Amazon is tell me best monitoring strategy is watch Netflix. If is up, they can able blame customer. If is down, they are fuck."
<html><body><b>Http/1.1 Service Unavailable</b></body> </html>
... then an empty console saying "loading" for the last 20 minutes. Then recently it upgraded to saying "Request limit exceeded." in place of the loading message (because hey, I'd refreshed the page four times over the course of 20 minutes).On the upside, their status page shows all green lights.
12:07 PM PDT We are experiencing elevated error rates with the EC2 Management Console.
They have standardized icons to represent various levels of issues (orange = perf issue, red = disruption). But they don't even use them. Instead they add [i] to the green icons to indicate perf issues (Amazon Elastic Compute Cloud - N. Virginia) and disruptions of service (Amazon Relational Database Service - N. Virginia).
Maybe this status page is controlled by marketing bozos who want to pretend the situation is not so bad.
EDIT: We are also using multi-AZ RDS, so either Amazon's claims for multi-AZ are bs, or their claims that this is only impacting a single zone is bs.
Interestingly, EBS will never return an I/O error up to the attached OS, which is likely a good decision as most OSes choke on disk errors. What this means, however, is that if something even get slow within EBS (let alone stuck), applications that are dependent on it will suffer. Most of these applications (such as databases) have connection/response timeouts for their clients, so while EBS might just be running slowly, a service like RDS will throw up connection errors instead of waiting even a bit more.
You can imagine the cascading errors that might result from such a situation (instance looks dead, start failover...etc)
If netflix is down, then it's something most companies who know how to design fail over can't cope with.
Other services, like Twilio, have come through several of these major problems with US-EAST generally unscathed while Netflix has had issues repeatedly.
Read their report from the major outage earlier this year, they start out by saying "elevated error rates", when many services were in fact down, and it wasn't until hours later they finally admitted to having an issue that affected more than just one availability zone.
From Forbes: ”We are investigating elevated errors rates for APIs in the US-EAST-1 (Northern Virginia) region, as well as connectivity issues to instances in a single availability zone.” By 11:49 EST, it reported that, ”Power has been restored to the impacted Availability Zone and we are working to bring impacted instances and volumes back online.” But by 12:20 EST the outage continued, “We are continuing to work to bring the instances and volumes back online. In addition, EC2 and EBS APIs are currently experiencing elevated error rates.” At 12:54 AM EST, AWS reported that “EC2 and EBS APIs are once again operating normally. We are continuing to recover impacted instances and volumes.”
A: fine A-: problems B+: servers are on fire
I really like Amazon as a company, use a lot of their services, but this is dishonest.
Except HN.
Cheers
Poisonous ideas spread as jokes. That is one of the ways they spread. A person thinking well about the issue wouldn't find the joke funny because it doesn't make sense. The joke relies on some poisonous, bad thinking to be understood. It has bad assumptions, and a bad way of looking at the world, built in.
It's akin to a man with athlete's foot deciding to remedy it by discharging a shotgun into his leg.
The Amazon thing in question is far more complicated, and far harder to understand.
Thinking they are "akin" is a mistake. It shows you're thinking about it wrong and failing to recognize how completely different they are.
One isn't going to confuse anyone or be misunderstood, the other will confuse most of the population and be misunderstood by most people.
One, if someone misunderstood, only involves one individual being an idiot. The other involves a large company being evil and thus can help feed conspiracy theories.
I'm not sure if you are aware of the difficulty of Amazon doing this. Suppose Jeff Bezos wants to do it. He can't simply order people to do it because they will refuse and leak it to the media and he'll look really bad and then he'll definitely have to make sure to try super hard for there to be no outages anytime soon.
Shooting yourself in the foot is stupid but easy. Doing this is stupid and essentially impossible. To think it's possible requires thinking that Amazon has a culture of unthinking obedience, or has an evil culture that all new hires are told about and don't leak to the media. Totally different.
Casually talking about impossible, evil conspiracies by big business, as if they are even possible, is a serious slander against those businesses, capitalism, and logic. Slandering a bunch of really good things -- especially ones that approximately 50% of US voters want to harm -- and then saying "it's just a joke, it's funny" is bad.
No one will believe its related and its certainly not slander to joke about it. Also you might want to leave the political opinions out of hacker news... there is no 50% of US who dislikes those things, they only have different ideas about how to support it.
I have to deal with a number of folks who will be overjoyed to read this news when their tech cartel vendor of choice forwards it this evening.
There's a huge contingent of currently endangered infrastructure folks (and vendors who feed off them) out there who throw a party every time AWS has a visible outage.
The perception that many people has is that somehow the 'cloud' is a magical up-time device that will save you money in droves.
The reality is that for many companies with mid-level traffic and aren't a start-up with a billion users, the 'old' style tech might very well be your best option.
Even if you're totally sold on the cloud, you can still have a requirement that things be transparent all the way down. AWS is one of the least transparent hosting options around.
If you're a customer of a regular colo, or even a managed hosting provider themselves based at a colo, it's pretty easy to dig into how the infrastructure is set up, identify areas where you need your own redundancy, etc. Essentially impossible within AWS -- there is no reason intelligible to me that ELB in multiple AZs should depend on EBS in a single AZ, but that's how they have it set up.
At the going rate they'll be AA members before long.
http://www.slideshare.net/twilio/highavailability-infrastruc...
http://www.twilio.com/engineering/2011/04/22/why-twilio-wasn...
It's strategy as opposed to how-to but the principles apply.
http://techblog.netflix.com/2011/07/netflix-simian-army.html
They've even released the "Chaos Monkey" open source: http://techblog.netflix.com/2012/07/chaos-monkey-released-in...
There's pretty much no way to architect around that one as an AWS user (apart from going fully multi-cloud, but "nobody" actually does that, at least at scale), and I'm kind of shocked that those bits of AWS are still not robust against "single AZ outages", given that they're involved in pretty much every one of these incidents and make them affect people on the entire cloud...
Pirate Bay might disagree with that sentence: http://torrentfreak.com/pirate-bay-moves-to-the-cloud-become...
But regardless it's not like all of EC2 went down just one or two AZs. So why couldn't traffic be migrated transprently to other AZs/regions ?
I would argue that none of the common full stack frameworks that startups use are fault tolerant enough for AWS. Most of them have multiple failure points that can quickly bring down entire apps.
After this issue is over I can give a longer answer. In short, we've just evacuated the affected zone and are mostly recovered.
And +1 for the slideshare page.
Their techblog is also worth following: http://techblog.netflix.com/
Since you've mostly recovered, how did your system do? Are there side-cases that Chaos Gorilla didn't touch?
Currently, Netflix uses a service called "Chaos Monkey" to simulate service failure. Basically, Chaos Monkey is a service that kills other services. We run this service because we want engineering teams to be used to a constant level of failure in the cloud. Services should automatically recover without any manual intervention. We don't however, simulate what happens when an entire AZ goes down and therefore we haven't engineered our systems to automatically deal with those sorts of failures. Internally we are having discussions about doing that and people are already starting to call this service "Chaos Gorilla"."
Heroku should offer a choice between N Virginia and Oregon hosting (I think they're almost comparable in price nowadays). That way people who want more uptime/reliability can choose Oregon. Sure it will be further from Europe (but then it will be closer to Asia) and people can make that choice on their own.
But basing an entire hosting service on N Virginia doesn't make sense anymore, considering the history of major downtime in that region.
http://www.forbes.com/sites/kellyclay/2012/10/22/amazon-aws-...
I can tolerate EC2/EBS going down but why on earth is ELB/Console always going down at the same time ?
I've been working with AWS since early 2006 when they first launched - I was lucky to be granted a VIP invite to try out EC2 before everyone else, and ended up launching the first public app on EC2. This might be the first time when frustration has overcome my love for these guys.
Also, the availability zones are disparate in terms of what they can support. A great number of my instances are in 1d because of unavailability in others.
Of course, as soon as I read the report that the issue was confined to one AZ, I looked to move my server over to another AZ. Oh yes, two were full and refused new instances, and then surprisingly, new requests for the other AZs never were received or operated on - and now the console is failing. It's a bit more than just "slow EBS" if that's what you were thinking.
- edit: said ec2 twice in the second sentence, corrected to say ebs.
Does anybody have any experience with migrating to Oregon or N. California in terms of speed and latency?
What the hell is happening here?
Just one multi-AZ RDS instance claimed the automatic failover. However, the 200+ alerts over the automatic failover due to internal DNS changes to point to the new master shows that things aren't as easily described by the RDS DB Events log.
Some instances reported high disk I/O (EBS backed) via the New Relic agent (the console still has some issues).
So far, this is what I see from my side.
2:20 PM PDT We've now restored performance for about half of the volumes that experienced issues. Instances that were attached to these recovered volumes are recovering. We're continuing to work on restoring availability and performance for the volumes that are still degraded.
We also want to add some detail around what customers using ELB may have experienced. Customers with ELBs running in only the affected Availability Zone may be experiencing elevated error rates and customers may not be able to create new ELBs in the affected Availability Zone. For customers with multi-AZ ELBs, traffic was shifted away from the affected Availability Zone early in this event and they should not be seeing impact at this time.
I'm normally an Amazon apologist when outages happen, but this is absolutely ridiculous.
http://www.reddit.com/comments/hg9oa/your_platform_is_on_aws...
... which you'll be able to click as soon as reddit is back up :)
@caseysoftware thanks for the links, we now have something to read and implement this week.
You have burstable network connections which by their nature, will have hard limits (you can't burst above 10Gbps on a 10Gbps pipe, for example; even assuming the host machine is connected to a 10Gbps port).
Burstable (meaning quite frankly, over-provisioned) disk and CPU resources.
And if any piece fails, you may well have downtime...
It is always surprising to me, that people feel that layering complexity upon complexity, will result in greater reliability.
- from a consistent average of 10ms over the last week
- to a new consistent average of ~2.5ms
beetwen 17:31 and 17:35 UTC. AWS started to report the current issue at 17:38 UTC. My app then experienced some intermittent issues (reported by newrelic pinging it every 30s). Don't know if it's related, could it be some sort of hw upgrade that went wrong ?
I did a push affecting my most used queries but that was one and a half hour sooner, at 15:56 UTC, so probably unrelated.
For most applications I think architecting EBS out should be straightforward - instance storage doesn't put a huge single-point-of-failure in the middle of your app if you're doing replication, failover, and backups properly.
And EBS seems to be the biggest component of the recent AWS failures upon which they've built a lot of their other systems.
Its high time AWS does something now.
I was just in the middle of booking a stay in a Palo Alto startup embassy for this week, too!
Deleted comment
Cloud hosting is not drastically different from any other type of service and is still vulnerable to the same problems.
That said, I don't think most people think using the cloud means that downtime is a thing of the past. I think the more attractive proposition is when hardware breaks, or meteors hit the datacenter, etc, it is their problem, not yours. You still have to deal with software-level operations, but hardware-level operations is effectively outsourced. The question is if you think you can do a better job than Amazon -- some companies think they can, most startups know they can't.
S3 probably fits the description of a cloud service. You send your data, and the service worries about making it redundant without your intervention. If data in NE USA is unavailable, the service will automatically serve you the data from somewhere else. You don't need to know how it works.
EC2 and some of these other building blocks, however, I would not consider to be cloud services. Merely tools for building out your own cloud services to other customers who then shouldn't have to think about failover and other such concerns.
If you know you are using a server that is physically located in a certain geographic location, it need not be represented by a cloud. It is a distinct point on the network.
This has to be a strong contender.
For many sites, a single server in a single zone (e.g., a non redundant server, an instance, a slice, a VM, whatever) is the right decision for ROI.
For many sites, the money spent on redundancy could be better spent on, say, Google Adwords, until they're big enough that a couple hours downtime has irreplaceable costs higher than the added costs of redundancy (dev, hosting, admin) for a year.
-The right to not have your traffic limited, and controlled by ISP
-The right to purchase a non DRM "file" and use it on your phone, computer, etc free of burned from some company
-Ability to install what ever you want on your $600+ device
*Edit for: formatting, and additional thought.