Solving the Mystery of Link Imbalance: A Metastable Failure State at Scale
code.facebook.com
code.facebook.com
Not just connections which traverse aggregated links. If you're load-balancing between multiple database replicas -- or between several S3 endpoints -- this sort of MRU connection pool will cause metastable load imbalances on the targets.
What I don't understand is why Facebook didn't simply fix their MRU pool: Switching from "most recent response received" to "most recent request sent" (out of the links which don't have a request already in progress, of course) would have flipped the effect from preferring overloaded links to avoiding them.
Without a bit more information I think I am just speculating and don't really know. Their solution doesn't really taste bad to me though.
Within their solution just had to convert the LIFO stack into an FIFO queue, keeping it simple.
EDIT: Also, this seems like a classic load balancing problem: Simply picking the least loaded connection would have been sufficient. The response time on each connection could be computed either explicitly (a running average/standard deviation of RPC finish times) or implicitly by checking the queue backlog at any time (the queue backlong being a first-order statistic).
One of these options makes things simpler. I'm always in favor of that.
I guess it depends how much minimizing the size of your connection pool matters.
That's probably the piece of that article that I'm going to remember for a long time.
This results in me exclaiming every 4-6 months or so that I wouldn't know how to create this particular bug on purpose if I wanted to, which I suppose doesn't make much sense as a complaint about a bug until you start thinking this way. Anyhow, it isn't perfect and you will sometimes be defeated by the sheer perversity of bugs and their behavior. (I also find these are the tiny ones, like OR instead of XOR or something equally simple and at times even one-character, that produce mind-blowing behavior off that one error. The architectural bugs tend to give way to this analysis much more readily.) But it's a useful tool much of the time.
I love reading about post-mortems like this, even if they're unlikely to happen at my startup, because the problem-solving techniques that get displayed tend to generalize to things of almost any size.
This is pretty neat, as the application can default to a round robin distribution and dynamically weigth the link selection key based on detected congestion/overload to shift the load to other links - since then I've often wished the TCP/IP world offered something similar.
Not a network engineer, but I had thought that TCP handles ordering itself? Packets can travel completely different routes from origin to destination, so one can't expect them to arrive in order. Destination's TCP stack can deal with it.
Again, IANANE, but ISTM that those who first designed the network introduced an untested complication with their LIFO setup. Nearly anything would have worked, including not pooling at all. They just chose something weird for the hell of it. Later FB needed more performance out of the system, and this harmful complication was hidden from sight.
In Linux, you can change the sensitivity to disordering by writing to /proc/sys/net/ipv4/tcp_reordering.
Out-of-order delivery causes TCP to shrink the congestion windows, which cuts throughput. Connection pools help here (in addition to their reduction in setup and teardown work), because they let the windows widen and stay open. We disable tcp_slow_start_after_idle to take advantage of this.
An insightful way of looking at a problem when the cause is not obvious.
It took a week or two to figure out why we couldn't transfer data but could connect, ssh, ping, etc.
I think my intuition there stems from having a lot of full-stack responsibility (so the idea that it's someone else's problem was never a luxury I could afford, since it was always my problem) and having cleaned up a number of other people's very large spaghetti-code disasters.
The other possibility is that I have no special intuition into the problem (this is far more likely) but did have a fresh set of eyes, while all the people who were working on the problem were so intimately familiar with it that they couldn't see the forest for the trees.
Well, sure, it's obvious when you're reading a prepared text that is carefully leading you up to that conclusion and has removed all extraneous information. Almost all bugs are shallow once you already know the answer.
Had there been some kind of negative feedback (in the electrical engineering sense of the word) and this still happened somehow it might be a lot more difficult to track down. But introducing negative feedback is what they did to solve the problem, so perhaps it wouldn't have still happened.
Thanks!
https://code.facebook.com/posts/360346274145943/introducing-...
Does the server side flow hash manipulation dependent on ip tuple manipulation? Only enabled for some hosts/devices? A bit disappointed, was hoping for clever dscp or mpls tag manipulation. If they rely on new streams with specific src ports its a lot less interesting, and less useful for long lived connections.