> Optimizing tail latency isn't about saving fleet cost.
Indirectly, it is. As the quote I replied to suggests, in order to combat tail latency services often run with surplus capacity. This is just a fundamental tradeoff between the two variables mentioned.
So, by improving the LB algo, they (and anyone, really) can reduce the surplus needed to meet any specific SLO.