Also keep in mind how percentiles compound when you have more than one service involved in serving a customer request. For example, let's say it takes 5 internal requests to serve an external customer request and each of those services measures latency SLAs at the 99th percentile. The customer request may only finish inside the SLA 95% of the time (99%^5)
Agreed, which is why the most surprising part of this article for me was this phrase: "In our main data center serving non-China traffic".
The transit latency to the one main datacenter in the (non-China) world is a lot more significant than the server time here. (Over 50 ms just to cross the US; compare with their server-side latency numbers of 95%ile < 5 ms, 99%ile < 50 ms.) If they're serious about latency, their deployment is holding them back much more than their choice of programming language or algorithm.
Or maybe the client is not the user's phone, but some other service running within the same datacenter?
Given that the service probably did not have a 100% success rate, and almost certainly had timeouts, the "max" would also likely be at the timeout.
You filter out the 504s, just the same way you would analyze it on a per-route basis for those metrics.
Quantifying things like service latency isn't a one-size fits all thing. Every service has its nuances and use cases that make it more meaningful to measure 99%, 99.9%, 99.99% or something else.
.... just don't measure average like I've seen naive junior devs do. Average latency is the worst of all metrics to use as it will include all the outliers at the very tippy top of the spectrum and basically render the metric meaningless.