Tail latency might matter more than you think
brooker.co.za
brooker.co.za
[0] https://www.infoq.com/presentations/latency-response-time/
Gil Tene also developed HdrHistogram, which has been ported to a bunch of languages, and is a great, lightweight way to collect accurate performance histograms: http://hdrhistogram.org/.
Support or Contact
Don't call me, I won't call you.I think an important idea here is that you should be trying to measure the experience of a user (or as close as possible). If there is a slow service somewhere in your stack, but has no impact on user experience, then who cares? Conversely, if users are complaining that the app feels sluggish, then it doesn't matter if all your graphs say that everything is OK.
I find it helpful to split up graphs/monitoring into two categories: 1) if these graphs look fine then the service is probably fine, and 2) if problems are being reported then these graphs might give an insight into why things are going wibbly. In general, we alert on the former and diagnose with the latter. Of course, its nigh on impossible to get perfect metrics that track actual user experience, but we've definitely found it worthwhile to try and get as close as possible to it.
---
Another fun problem with using summary statistics is they can easily "lie" if the API can do a variable amount of work. For example, if you have a "get updates API" that is called regularly to see updates since the last call, then you end up with two "modes": 1) small amount of time between calls and so super fast and 2) a large amount of time between calls and so is slow. Now, in any given time period the vast majority of the calls are going to be super quick, but every user will hit the slow case the first time they open the app for the first time that day. This results in summary statistics that all but ignore those slow API calls when opening the app.
Sure, it gets messier, and definitely less visually appealing, but the reaction by others has uniformly been "Did we have this data available all along and just never showed it?!"
It's also worth mentioning that even in a system where technically tail latencies aren't a big problem, psychologically they are. If you visit a site 20 times and just one of those are slow, you're likely to associate it mentally with "slow site" rather than "fast site".
In every case it has been useful and actionable.
90th percentile is the fast path. You notice overall effects of growth, load, regressions, etc. 99th percentile (or sometimes 99.9th) is often the actual SLA/SLO so it should be on the plot. Max, as you point out, is necessary to actually see the long tail.
The other thing that's important is to make sure the timeseries and plotting software isn't averaging the max value; it's easy to miss spikes of latency because someone thought the graphs looked nicer with avg(1 min) vs max(1 min) or a tsdb is automatically rewindowing/aggregating with the wrong function. To a lesser extent even percentile buckets can be deceiving if they're too big. 99th percentile with 5-min buckets can miss significant but brief (~3 second) degradations.
What a strange way to put it. No one wants to see the long tail. The long tail is the means, not the end.
To a lesser extent even percentile buckets can be deceiving if they're too big. 99th percentile with 5-min buckets can miss significant but brief (~3 second) degradations.
Now that's the actual reason to care about max. Otherwise you have values (and often negative user experiences) getting lost in the aggregation.
Really? I find percentiles to be more informative. Looking at the maximum is like looking at an error log, it is basically throwing out all performance data except data from a tiny slice of time. A bad maximum shows you that a service failed at least once, but you oftenalready knew or expected that. A bad 90th percentile tells you that a lot of users are experiencing poor performance.
That said, I do look at percentiles in addition to the max. But I find the 99th, 99.9th, and the max together tells me a lot more about the system performance than the uninformative stuff at the 90th and below.
It's a good point, even though it cuts both ways: given some assumptions about the tail behaviour of latencies, the 20 minute extreme event is a treasure trove for estimating the probabilities of smaller tail events.
I tend to look at them per request, per user, per iteration, and so on, to control for that effect.
For the record, I bundle all my JS/CSS assets in one file.
Having circuit breakers and carefully tuning timeouts can help.
Can help, for sure. Can also turn a 'slow' into 'unavailable', which could be what you want, or could be a disaster.
This is the #1 thing that people get wrong. It's something that otherwise smart software engineers get wrong because they don't have enough of a background in data analytics.
The problem has two parts. One part of the problem is that once you reduce an observation to summary statistics, you can't go back. The other part of a problem is that web services usually generate too much data if you don't summarize.
Typically you also want to record backend information. One HTTP request at the front end might be twenty or a hundred backend requests.
Just one example that comes to mind: suppose you want average (mean) latency. So you record, for every minute, the average latency for requests during that minute. Now you want to calculate the average latency for an entire day… but it’s impossible. You can calculate the daily average of the per-minute averages, but you’re averaging over minutes, rather than averaging over requests.
This is just a simple example, but you can see how a seemingly innocuous decision sabotaged your ability to run the query you want.
Maybe my advice is “learn calculus”.
So, well meaning engineers will collect percentiles on individual components in the system, with the idea that if you want percentiles at the end, just collect percentiles everywhere. You’re left with garbage data that can’t be aggregated in a way that makes any sense.
Histograms can be added and aggregated. You can always estimate the mean from a histogram.
So if you compute a histogram of latency values every minute, you can add them up for successive minutes to get the distribution over larger intervals.
But the problem with averaging averages is that it makes no sense mathematically. The average (well, arithmetic mean, to be technical) has a well-defined meaning and interpretation in terms of the original data points. When you start averaging averages (or dear Lord percentiles!) you lose the last glimmer of interoperability you had.
The bigger issue here, aside from the fact that you may be taking a simple average when you need to be taking a weighted average to aggregate your individual point estimates, as others have pointed out, is that you can estimate a sufficiently symmetric Poisson distribution with a normal distribution, but response latency is not symmetric, specifically because time travel isn't possible and you can't ever get negative response times to balance out the tail events on the right side of the distribution. Tail events with respect to latency can only be bad, not good. They're a bet with a hard cap on winnings but potentially unbounded losses. As such, you really need to consider what the worst case is and the expected frequency of that worst case.
For one great deep dive on the topic, check out https://drkp.net/papers/latency-socc14.pdf
I suspect, naively, that latency purely from network appliance queueing probably gives a mixture of Poissons in a straight line bus topology. But add in arbitrary routing and the much more complicated schedulers of the servers themselves, and you get no well-behaved distribution at all.
It's especially interesting to compare the SLAs given by data center providers that tend to always use 99th or even 99.9th percentile worst case performance, with application metrics using arithmetic averages.
Human height?
- Minute 1: 10 requests, 5 with 10ms latency, 5 with 500ms latency; avg latency 255ms
- Minute 2: 1000 requests, 900 with 10ms latency, 100 with 500ms latency; avg latency 59ms
Average of averages: 157ms
Average latency: 97ms
You can make them arbitrarily different.
> (255ms + 59ms) / 2 = 157ms
But
> ((10 * 255ms) + (1000 * 59ms)) / 1010 = 60.94ms
Which is the same answer I get as averaging all the samples without bucketing. (I can't get 97ms though)
So if your response is, “can’t you just collect the correct data in the first place?” Well, that’s the hard part.
If all you have is 50 averages, there is no way to combine them to get the average of the underlying data, because you don't have the original "length(x)".
To go to the next level I would suggest also emitting the sum and sum of squares, which will make your summary statistic far more useful.
And choose the right quantization for histograms/cdfs.
> To go to the next level I would suggest also emitting the sum and sum of squares, which will make your summary statistic far more useful.
Histograms should also give you the sum. I haven't seen sum of squares in many monitoring setups; is this for underlying metrics that are differences/errors?
This allows aggregation for efficiency, usually either by partitioning the data by short time windows (1 second or more usually) and storing measures over each duration or by storing a sampled subset of original values, both of which can later be queried to produce measures over arbitrary periods of time.
A few measures can be supported exactly with pre-aggregation (min, max, mean, variance), some can be supported with bounded error (quantiles) or statistical error (pretty much any measure over sampled values).
[1] https://observablehq.com/@oavdeev/parallel-task-latency-calc...
When you see a beach ball or other indicator of delay on your computer you are quite likely to be experiencing tail latency.
Using streaming/event buses and getting rid of the chain of calls in front of the initial request definitely has upsides. Syncing data all over the place, with consumers being forced to interpret that data from outside of their domain has large downsides too though.