The more requests you have, the higher the chance one of them hits a tail. So the overall latency a user sees is largely dependent on a) number of requests b) tail latency of each event.
This method improves the tail latency for ALL supported services, in a generic way. That's multiplicative impact across all services, from a user perspective.
Presumably, the number of requests is harder to reduce if they're all required for the business.
> Prequal has dramatically decreased tail latency, error rates, and resource
> use, enabling YouTube and other production systems at Google to run at much
> higher utilization.Indirectly, it is. As the quote I replied to suggests, in order to combat tail latency services often run with surplus capacity. This is just a fundamental tradeoff between the two variables mentioned.
So, by improving the LB algo, they (and anyone, really) can reduce the surplus needed to meet any specific SLO.