Twitter Completely Down
isup.me
isup.me
I'm fairly certain that the higher-ups in Twitter weren't told "We have pretty good failover protection, but there is a small risk of catastrophic failure where everything will go completely down." Whoever was in charge of disaster recovery obviously didn't really understand the risk.
Just like the recent outages of Heroku and EC2, and just like the financial crisis of 2008 which was laughably called a "16-sigma event", it seems pretty clear that the actual assessment of risk is pretty poor. The way that Heroku failed, where invalid data in a stream caused failure, and the way that EC2 failed, where a single misconfigured device caused widespread failure, just shows that the entire area of risk management is still in its infancy. My employer went down globally for an entire day because of an electrical grid problem, and the diesel generators didn't failover properly, because of a misconfiguration.
You would think after decades that there would be a better analysis and higher-quality "best practices", but it still appears to be rather immature at this stage. Is this because the assessment of risk at a company is left to people that don't understand risk, and that there is an opportunity for "consultants" who understand this, kind of like security consultants?
Is this really a valid conclusion to come to at this point? I expect downtime in any service I operate. It's just how the world works. Does that mean I don't understand the risks and am misleading the board?
(expanding my reply) No risk assessment in the world will stop a cage monkey from tripping over a pile of 1Us and falling onto the big red button. Figure out what your pain threshold is and live with it.
Just finding it took days to sort out, apparently.
statistical distributions that include a risk aversion parameter exist precisely to model this kind of problem (unfortunately their names have slipped my mind, otherwise I'd provide links)
But also to the thread parent's comment about Twitter execs not realizing there was a minor risk of catastrophic failure: Executives only care that their money-making baby keeps running. In the past i've seen execs demand that an engineer call them at 3AM if the production site goes down for more than 5 minutes... even though that call is pointless. I think they just assume there's no point in getting involved with the plan because the plan will never be perfect, but at least they can be aware of a problem so they can cover their asses and tell a higher-up that it's being worked on. At the end of the day, even the guys at the top don't really give a shit about the product, they just care about their paycheck.
The point is that at least in the case of Heroku and EC2 (I'm not sure what caused Twitter's outage yet), the causes of failures weren't something like a "16-sigma" event like a plane hitting an electrical tower (a tragic event that happened in Palo Alto a couple of years ago). They were things like insufficiently tested software and processes, and misconfigured devices. These things do not add cost, except maybe incrementally more man-hours in terms of testing and auditing of configurations. They are not million-dollar diesel units that require permits, etc.
My point is that if a device misconfiguration can take down EC2, or a single bad data in their data stream can cause massive failure, it means that the entire system is much more fragile than its been sold to everyone. If they didn't realize this was the case, it means that the risk of failure was a lot higher than they had assessed.
That's not necessarily true. People don't die when twitter is down, and whatever twitter's business model actually is, I am not even sure there is a monetary penalty to them being down (unlike, say, Amazon being down which results in lost orders). They may have made the calculation that it was not cost effective engineering-wise to chase that extra 0.001% of reliability.
[Edit: Pedantry shield: Ok, ok, should have said people don't die because twitter is down. Obviously people are dying all the time, and some will indeed expire while twitter is down].
There is no such thing as bad publicity.
And here I was hoping we could just take down Twitter and live forever...
Any assessment of risk entails certain larger assumptions about the world, many of which often turn out to be mere guesses.
Consider all the prices that are set to their current levels b/c nobody expects the collapse of the US political system to occur. Yet there is a nonzero probability that it will occur.
On one hand this seems like an absurd example, yet it exemplifies the kind of blind spot we are prone to when assessing risk. We generally address all the risks we can directly control, then classify the rest as "systemic" which essentially means that we are not able to compute them so we're going to ignore them.
Yet many systems which we assume to be stable or predictable (governments, companies, markets, weather patterns, social trends, etc.) have unexpected aberrations now and then which can have very significant consequences. Since these tend to impact most companies equally, the market will converge on an equilibrium where no firms do anything to hedge against these things.
Do you want to pay extra bank fees so that your bank can hedge against the collapse of the US currency for your checking account? Probably not. Do you want to triple your hosting costs to hedge against a massive US power grid failure? Probably not. The same applies to asteroid risk and sudden ice age risk.
On the other hand, if you have lots of money saved, you may wish to hedge against the collapse of one currency or another, and if your business would end if you suffered a few hours of downtime, you might want to invest in massive amounts of redundancy.
Every morning when we all commute to work we risk death. Some exposure to systemic risk is considered acceptable, and part of the character of any person or business is the kind of risk exposure we tolerate day to day. A doctor working in an AIDS clinic risks needle sticks and HIV, a startup doubling its users each month risks downtime but also risks a cash flow crisis.
It is likely that what you mean by "properly" is impossible. At large enough scales, what you end up with is a Gaussian distribution of errors in accordance with the Central Limit Theorem... except that there's a Black Swan spike in the low-probability, high-consequence events, and you basically can't spend enough money to ever get rid of them. Ever. Even if you try, you just end up piling equipment and people and procedures which will, themselves, create the black swan when they fail.
I think you're trying to imply that if only they'd understood better, this could absolutely have been prevented. No. Some specific action would probably have been able to avert this but you simply don't have a 100% chance of calling those actions in advance, no matter how good you are.
The state space of these systems is incomprehensibly enormous and there is no feasible way in which you can get all the failures out of it, neither in theory nor in practice.
Living in terror of the absolute certainty of eventual failure is left as an exercise for the reader.
It not only occurs in the technology industry, but also even in things like financial risk analysis. For example, people could mitigate the risk of a bond defaulting by buying a credit default swap. However, most people failed to assess the risk of their counter-party going belly up, like AIG or Lehman. This failure in risk assessment is in large part why the financial crisis was so widespread.
So, yes, many if not most people are very poor risk assessors.
Then again, this might be rather a capacity (or the lack thereof) of the organization within the risk is assessed. This is what the software engineering quip "Most problems are people problems" means, I think. In some environments it is hard to bring up the unlikely, catastrophic scenarions without being seen as overly pessimistic and somehow not enough subscribed to the success of the undertaking as a whole. So you assess risk success overly optimistic in order to further your career, not to assess risk accurately.
If they'd had a couple of gennies on the rooftops, they ... well, they wouldn't have been fine but they'd have had a fighting chance to keep the scrammed reactors from melting down. Or if they'd had a higher sea wall (like Onagawa) they'd have been fine.
So: not one, not two, but four power systems failed in order to result in the meltdowns -- and one of them would have worked if they had been located slightly differently.
(Otherwise, your point about people being poor risk assessors is spot-on. And worse: even if some people are acutely conscious of risk, once decision-making responsibility devolves to a committee, the risk-aware folks may be overruled by those who Just Don't See The Problem.)
My main beef with their failure handling is actually this: you need to be able to face a situation where _all_ you smart emergency systems fail. In the case of an NPP this can mean an almost global environmental crisis, and need relocate millions of people, making hundreds of square kilometers uninhabitable, etc. In that case you don't really want to rely on five generators on some roof. Which may or may not work on that day.
And this is not something I make up here now. I remember discussing nuclear safety in high school, and the bottom line was: NPPs are ok, since they become uncritical when _everything_ fails, because the moderator rods slide down into the reactor vessel.
But after Fukushima I read, that actually the situation there, with that specific model, is different, unfortunately. Tough luck. Because that model still needs some cooling because the fully moderated reactor still produces 1% it's total energy, an that is enough to bring the reactor into an 'undefined' state, iirc. And it is easy to imagine what that means for a station that has just been struck by an earthquake anyway.
My whole point is: your comment makes it appear as if the security layers actually were plenty, and I would (respectfully, of course) disagree with that. I think it was poor.
That point is important: if NPPs aren't build safely, what is then built safely? My guess is: nothing.
So what to do? Design for failure. (Politically, technically, economically, can be applied everywhere.)
End Of Rant :)
Yup.
On a similar note, the French response to Fukushima Daichi is rather interesting (France relies on nuclear generation for over 80% of its electricity):
http://www.nature.com/nature/journal/v481/n7380/full/481113a...
"The ASN has also come up with an elegant technical solution to get around the (universal) dilemma of how to protect a plant from external threats, such as natural disasters. The report recommends that all reactors, irrespective of their perceived vulnerability, should add a 'hard core' layer of safety systems, with control rooms, generators and pumps housed in bunkers able to withstand physical threats far beyond those that the plants themselves are designed to resist."
(And a mobile emergency force who can move in and stabilize a reactor after an unforseen catastrophic disaster that kills everyone on-site and destroys most of the safety systems.)
In other words, they now expect unpredictable Bad Things to happen and are trying to build a flexible framework for dealing with it, rather than simply relying on procedures for addressing the known problems.
Totally agree with you here - Until recently, no nuclear power plan was designed such that it could survive failure of all their emergency systems. Hopefully, with the negative repercussions of Fukushima on the industry, engineers are rethinking their approach to Nuclear Power.
http://www.theregister.co.uk/2001/09/10/its_bofh_disaster_re...
In almost every disaster situation I've been a part of, the UPSes have failed. Almost every time.
Disaster recovery has a horrible track record.
http://www.schneier.com/blog/archives/2009/08/risk_intuition...
Googling "Schneier risk" gets you lots and lots of reading material.
The only question here is whether Arrington is going to write another Amateur Hour post about this.
http://techcrunch.com/2008/04/23/amateur-hour-over-at-twitte...
One hour RTO and RPO will cost way more than 24 hour recovery. Edit... and the business managers decide how much they wish to spend on DR. It's a trade-off and anyone who has ever done it, understands that.
Risk management consultants already exist! Many companies just choose to assess it internally, or the consultants themselves are inexperienced (often lacking practical experience).
Disaster Recovery is a reactive approach. It's what you do to get things back up AFTER a system or site has failed.
Business Continuity is a proactive approach. It's what you do to ensure that your critical services will remain viable whenever disaster occurs.
In the cases of Heroku, Amazon, Twitter, and many more, their Disaster Recovery strategies have been successful. The fact that they came back online without major data loss is proof of that. Their business continuity strategies, however, have been found wanting.
Obviously, this does you no good when something like an API service goes down, but that's issue is irrelevant to the question that was posed. Having some static file to test against whether it is accurate or not will instantly tell you whether it is your code or the third party that is breaking.
I have a suite here that takes about 3 minutes to run from scratch, but just over 1 second as unit tests.
Sure, being upset/getting angry just because of a little bit of Twitter downtime is stupid, but that doesn't take away from the fact that one of the biggest and most important discussion and communication channels the web has is completely down.
Both Twitter going down for an hour and a plant that assembles cars going down for an hour aren't that big of a deal in the huge scheme of things, but for people who are intimately connected (either work there, know someone who does, are emotionally connected to the product in some fashion, etc), it feels a lot bigger than it is.
You know, learn from others mistakes and all.
The single most important part of the internet is the immediate availability of news (to me, anyway). I've never heard anyone complain about that before; why do you think it's not worth knowing and talking about events as they happen? 'Twitter is down' isn't noise. Years and years ago there were stories about fire departments (SF I think?) that started using Twitter to send out fire notices. Here in Montreal, the police tweet very quickly and accurately about our daily student protests.
Twitter is extremely important to a huge number of people, and when it goes down, a site like HN definitely should be talking about it. It's big news and it's almost exclusively relevant while it's happening.
I love the internet.
Think about these two scenarios:
1. During that one hour you spend your time focussed on talking about this event as it's happening, speculating, having an emotional response (because you can't access something you want to and find a group of people experiencing the same thing and all get together to experience the frustration).
2. Tomorrow you read a story that says "Twitter was down for one hour yesterday" with some detail about what happened.
I believe that the latter is preferable. It's more efficient, less emotional and more useful. The former is the same as watching some 'Breaking News' event while is happening.
Now imagine that the one hour of downtime happened when you were asleep. You've missed nothing.
There are two scenarios where this news is important: if your business depends on Twitter, and if you are trying to assess the reliability of Twitter. The latter can be achieved by #2 above, only the former needs real-time updates and that doesn't mean general news reporting just your own monitoring.
The 24 hr news cycle ruined TV news, and SEO has made Internet news worse (first links win).
Think of "twitter being down" as Silicon Valleys equivalent to hollywoods "Lindsey Lohan is drunk in jail again"..
The tech companies, their founders staff and services are our pop-culture to gossip about.
This is Hacker News, not Globe for Hackers. Seriously, tabloid-esque coverage of tech companies adds nothing of value at all.
the word has been appropiated by those who are needy like that: you don't call yourself hacker, just like you don't call yourself saint. only mediocre people to whom it never applied and never will would do that. the end.
Twitter is infrastructure for us in media. It's well worth discussion.
also, "us in the media", what kind of whore talk is that even? twitter is correctly referrerd to in w3c docs as medium preventing intelligent discussion. so it was down? GOOD. people are inconvenienced? even better! it cannot possibly have hit anyone or anything that was worth fuck all.
I'm not sure why you suggest RSS is somehow synonymous with Twitter, but I will say that in addition to RSS buttons almost every major media company on the planet has a Twitter button on its article page (I work on one, which is why I say "us in media"). Because many sites don't do proper async JS when it comes to social buttons, an outage on socnets can be crippling
Here's an example: http://techcrunch.com/2012/06/01/facebook-outage-affects-oth...
These posts are so incredibly annoying. We can see if service x is down for ourselves. That isn't news. I could maybe accept these stories if the link on the front page was to a blog post stating that not only is service x down but why it went down for sure plus an added lesson we can learn from it. Short of that it's become an easy way for people to build up a trillion karma points. And if you want to tell me you don't care about karma then you're either lying or you have none. Enough with this crap. We'll find out ourselves but most of us won't actually because we have lives and by the time we go online to check our favorite wank-off site it'll probably be back up again like the past fifty times I've seen a story about Heroku/AWS/Twitter being down.
Yes, they had a good idea and executed very well, but as I see it, Twitter is nothing more than your run-of-the-mill 4chan board. I still don't understand the draw, but then again, there's a lot of facets of modern society that I simply have no explanation for (reality tv?!?) and have been better off not worrying about it further.
I'm still skeptical and I have used it.
Of course, that woudl imply that the act of removing credibility from something would be "decrediblizaiton" or "decredencing".
Did you understand what the GP meant? Yes. Did you need to point out to the world your complete mastery of the English language? No.
The GP's point was that Twitter is useful, despite it's relatively low score on "innovative new technology" scale.
Something the vast majority of people couldn't care less about. Twitter has, in effect, created a new messaging protocol, in the broadest sense of it. It's accessible on your computer, your phone, even your TV if you try hard enough. It's integrated with hundreds of apps and sites. Technically speaking it isn't doing anything particularly amazing (although the sheer scale they deal with is), but that's not really the point.
Note the irony of this comment existing in a post about the service not being accessible, anywhere. Which wouldn't be the case had it been a protocol.
An IRC-styled Twitter would need some sort of synchronization service to handle twitsplits.
What is great about Twitter is that it allows both the sender and the receiver to choose if they want to interact over the web, through an app, via SMS or even email (receiving DMs). Take the lowest common denominator across all platforms (no subejct, max 160 chars in SMS) and sit in the middle as an abstraction.
Unfortunately there are no details, it just says "there was a cascading bug in one of our infrastructure components".
No news at the status site either, that beats the purpose of having a dedicated status site.
Users may be experiencing issues accessing Twitter. Our engineers are currently working to resolve the issue.
It was unavailable to many people, not down.
That's called understatement.
Yep. Even the subthread from the person complaining that this isn't newsworthy.
Edit: Nope. Just slow. My tweet appeared.
As a workaround for those systems, add s.twitter.com, search.twitter.com and api.twitter.com in your /etc/hosts file that map back to 127.0.0.1.
This obviously breaks Twitter integration, but it also makes sure page loads don't explode when waiting for remote resources.