Exponential Backoff and Jitter (2015)
aws.amazon.com
aws.amazon.com
One day one of our backend services became temporarily unavailable and caused millions (probably in 10s of millions at this point) of connected web clients to go into jitter-less retry loops. The retry requests were arriving in coordinated waves and their amplitude was so high that we had trouble bringing the service back up, causing further disruptions. Once we had dug ourselves out, the first thing we did was to make sure all retry loops had jitter.
After you published a message to that topic all 1M of them were closing a long polling HTTP request and immediately reconnecting. Which was causing a thundering herd problem [1], but I only learned this term much much later. The obvious solution was to add randomized delays to reconnects.
Of course that was probably the easiest part of scaling that system.
I'm interpreting that congestion may be the result of the client expectation of sending into a queue based on the maximum possible availability of a channel, whereas if you add randomness to the percieved availability of the channel across all clients, all clients will then send based on the perception/calculation of the mean availability, which is always less than the max as they all converge/revert to that mean.
(crazily, it could imply if we halved speed limits but didn't really enforce them, people would mentally reserve twice the time duration for a given trip, and therefore only choose discretionary travel based on that doubled time commitment, reducing the number of people using the roads together at a given time, while the ones on the road could travel double the percieved speed limit)
[1] https://cs162.org/static/readings/jacobson-congestion.pdf [2] https://www.i-programmer.info/babbages-bag/398-ethernet.html... [3] https://www.cs.utexas.edu/users/lam/NRL/backoff.html
It is absolutely full of really interesting topics that don't get covered frequently enough other than in footnotes or in overly simplified hand-waivey explanations. AWS has done a great job in trying to demystify some complex topics.
Well, those things go on and off at random but then the mobile operator fixes a huge network outage and boom, several million devices go online at once.
One common real world situation where this comes into play are things like thundering herd scenarios. Rather than trying to cope with huge bursts of traffic, we want to spread the load out in time.
Exponential Backoff and Jitter - https://news.ycombinator.com/item?id=9206090 - March 2015 (15 comments)