Facebook App Outage Postmortem
developers.facebook.com
developers.facebook.com
Maybe we (the IT community) need a framework for incident reports or postmortems, or just use Google's as a model?
[1] http://googledevelopers.blogspot.ca/2013/05/google-api-infra...
It doesn't seem crazy to me that Facebook's publicly facing summary of this is as casual as this seems to be. They owned up to breaking their platform and indicate they're taking measures to not do it again. But if the person who's internally accountable for analyzing this and preventing recurrence told me "we're building better tools" without any specifics about those tools, who's got accountability, or the timeline they anticipate putting those in place, I'd say they should pack their bags, so I bet there's a more detailed plan internally. I'm also not a facebook app developer, though, and if I had any revenue depending on not being shut down like this, I might be more frustrated with this either a) poor level of transparency (giving them the benefit of the doubt) or b) poor depth of analysis.
I used to work at Facebook, and this is most definitely not the internal audit. A lot of Facebook engineers are former Googlers, and bring a lot of the culture and practices with them. You can rest assured that people are hunkering down in a conference room as we speak.
That said, Google's postmortems are a thing of awe and distributed widely within the company.
For example, I haven't been able to view friends' profile pages on FB for days now... blank pages in any browser on any machine as long as I'm signed in. I'm sure they'll notice they broke it sometime in the next month and fix it.
Unlike, say, every single action actually described in the text...
Seriously though, how on earth do you even get the idea to initiate such an operation without a) testing it against realistic data, and b) doing a dry run against the live data first before you decide to pull the trigger and start terminating applications?
And to top it all off, there was no roll back scenario and the restore process was buggy, which means it wasn't tested properly either.
I'm all for "move fast and break things", but this is a complete joke.
The organizational chaos that leads talented engineers to operate this way must be a nightmare.