Maybe we (the IT community) need a framework for incident reports or postmortems, or just use Google's as a model?
[1] http://googledevelopers.blogspot.ca/2013/05/google-api-infra...
Maybe we (the IT community) need a framework for incident reports or postmortems, or just use Google's as a model?
[1] http://googledevelopers.blogspot.ca/2013/05/google-api-infra...
It doesn't seem crazy to me that Facebook's publicly facing summary of this is as casual as this seems to be. They owned up to breaking their platform and indicate they're taking measures to not do it again. But if the person who's internally accountable for analyzing this and preventing recurrence told me "we're building better tools" without any specifics about those tools, who's got accountability, or the timeline they anticipate putting those in place, I'd say they should pack their bags, so I bet there's a more detailed plan internally. I'm also not a facebook app developer, though, and if I had any revenue depending on not being shut down like this, I might be more frustrated with this either a) poor level of transparency (giving them the benefit of the doubt) or b) poor depth of analysis.
I used to work at Facebook, and this is most definitely not the internal audit. A lot of Facebook engineers are former Googlers, and bring a lot of the culture and practices with them. You can rest assured that people are hunkering down in a conference room as we speak.
That said, Google's postmortems are a thing of awe and distributed widely within the company.