Details on today's Facebook outage
facebook.com
facebook.com
The thundering herd problem occurs when a large number of processes waiting for an event are awoken when that event occurs, but only one process is able to proceed at a time. After the processes wake up, they all demand the resource and a decision must be made as to which process can continue. After the decision is made the remaining processes are put back to sleep, only to wake up again to request access to the resource.
This occurs repeatedly, until there are no more processes to be woken up. Because all the processes use system resources upon waking, it is more efficient if only one process is woken up at a time.
This may render the computer unusable, but it can also be used as a technique if there is no other way to decide which process should continue (for example when programming with semaphores).
Though the phrase is mostly used in computer science, it could be an abstraction of the observation seen when cattle are released from a shed or when wildebeest are crossing the Mara River. In both instances, the movement is suboptimal.
Since those systems are doing very little IO, configuring connection pools to start with much more connections than the DB has CPUs and to add more if the connections are busy (i.e. DB is getting slower, probably because it is highly loaded), is guaranteed to cause resource contention on the DB and escalate the issue in case something goes wrong.
Its a total waste and yet the most common configuration in the world.
Are there any architectures/patterns/methods that can help make it easier to find the source of performance issues?
Since we encounter this on a regular basis we have built a few different systems to gracefully handle them.
Unfortunately, the event today was not just a thundering herd because the value never converged. All clients who fetched the value from a db thought it was invalid and forced a re-fetch.
The strangest thing to me was that the clients had the permission to essentially cause a race condition across the whole site. From what I understand, the client ran an API call which either forced a re-fetch of a key which apparently only needed to be fetched once (and thus theoretically could have been staged in advance using a database application not running on the frontend site which could update the cache and prevent a database fetch, or just update the database, whichever was more necessary), or failing the database connection due to the aforementioned thundering herd of QPS it also triggered a re-fetch from the db (which again could have been prevented by a db app pre-loading the new value). So, if my outsider's idea of how Facebook's code works is accurate, this could have been prevented if either the cache/database was "pre-fetched" in the background not using client API calls, or if client API calls simply weren't allowed to all modify [and read] from a single key in the database. The latter point seems less likely than the former, but possible.
(Sorry if i'm speaking out of line or as an uneducated FB user, but according to these comments and RJ's breakdown this is how it appears to me)
I guess at Facebook's scale you have to build in fallbacks but this is a reminder that you can easily do more harm than good.
Of course, doing whole-stack testing in a meaningful way at Facebook's scale is a very tricky thing.
Which makes integration testing all the more critical.
Stick with Mysql
* * * * If it aint broke, dont fix it!
Me too Melissa, and it's out there in the media that a group of hackers caused the problem, is this true, Mr Robert Johnson?
PLEASE !!!! WHAT CAN YOU DO TO HELP THIS FROM EVER HAPPENING AGAIN???????????????? PLEASE!!!!!!!!!!!!!!! CANDY
Kip da updates comin'
Did anyone get a message like I did about someone trying to access your account from another state?"Marvin, the server was like a dog chasing its tail...it kept going in circles, but never caught it. They basically had to hold the tail for the dog so he could bite a flea on it. :) LOL"
"You are using an incompatible web browser.
Sorry, we're not cool enough to support your browser. Please keep it real with one of the following browsers:"
This is why I don't use Facebook. It's not 1990 anymore. You don't need User-Agent sniffing.
Angee Lening Marvin:
the server was like a dog chasing its tail...
it kept going in circles, but never caught it.
They basically had to hold the tail for the dog so he
could bite a flea on it. :) LOLInteresting article, terrible comments.
That's just a guess. But from their post, it seems reasonable.
The more likely explanation is that someone on 4chan noticed FB was down -- and then started talking about DOSing it.
After facebook's outage started, there were like 9128 threads in /b/ with trolls claiming it was their success.