The Most Misleading Measure of Response Time: Average
blog.optimizely.com
blog.optimizely.com
If you need to reduce a distribution to a single number, the most informative number is going to be the mean.
I understand their point about the 99th percentile, but consider that it's possible to improve the 99th percentile measure, while increasing the mean and degrading the performance of all but 1% of the users.
The real issue is reducing a distribution to one number.
For example, most people intuitively perceive the fact that "the majority of drivers consider themselves above average" as human stupidity, however mathematically it makes perfect sense if average == arithmetic mean.
If you are not convinced, consider Bill Gates walking into a room full of college students - suddenly, almost everybody becomes below-average wealthy (if average == mean).
To me, it's not clear that there is a "most appropriate" number. All mappings of a distribution into a number will throw information away - the trick is to pick a number that keeps the information you want. And that seems to depend on the context.
I intuitively perceive it as good drivers being clustered together near the top, with a few outstanding lousy drivers near the bottom.
If you are not convinced, consider Bill Gates walking into a room full of college students - suddenly, almost everybody becomes below-average wealthy (if average == mean).
That's true in my country even if Bill Gates doesn't walk in. Hell, perhaps it's true even in the US.
No such rule of thumb is ever going to work, you always need to consider context.
In the case where disutility is non-linear with response time, especially if there is a cliff below which differences are irrelevant, an improvement in the response time to the worst decile may well be worth degraded performance for a majority of users. The most useful single statistic in that case could be, not mean, median, or mode, but the percentage of users who fall above the "unacceptable" delay threshold.
Not necessarily. If you want an general-purpose summary statistic that's easy to interpret, for many distributions, the median is a better statistic than the mean. (The pathological case is the Cauchy distribution, which looks like a normal distribution with heavier tails, but has a median but no mean, and both any individual data point and the sample mean approximate the median of the true distribution equally well.)
But yes, unless you have an a priori reason to think that the response is normal, it is worthwhile to look at quantiles, a histogram, or a kernel smoothed density estimate of the data as opposed to a single statistic.
But the key takeaway is to consider distributions and to consider confounding factors.
I've complained about this before; use maximum frame delay [0] instead of FPS when measuring UI responsiveness.
[0] The maximum time elapsed between any two sequential frames during your test.
Understanding where your users are on the curve is probably more interesting than a single number. Worrying about that last 1% really only makes a meaningful difference if your user base is huge enough that fixing something for 1% of your users can jump revenue by a significant multiple.
Mentally, I try and think of the 80 or 90% of users with a similar experience, needs, etc. and make it better for them. In this case, speed is good for everybody, but I care very little about the needs of that last 1% if your customers are all paying the same. No sense in putting the needs of a small number of users in front of the needs of a much larger set of users.
Anyway, I'm really glad that they've improved the load times for their snippet, because this issue is always a genuine concern that needs resolving.
Want to do it yourself? This talk by Etsy a few weeks ago has some detail on how they did a similar thing:
http://www.slideshare.net/marcusbarczak/integrating-multiple...
Some links at the end of the talk. Infrastructure wise, I think you have to be prepared to pay for some expensive DNS before this kind of thing is viable.
It is only about 13 pages, making it a quick but very informative read. I highly recommend it for anyone trying to measure performance, throughput, response time, efficiency, skew and load.
Direct link: http://pages.optimizely.com/CDNBalancingWhitepaper_GeneralLP...
My theory is that they're using the Dynect CDN manager. This is a tool created by Dyn to let companies use their DNS system to pick which CDNs are used where, and to create round robin rules that are percentage and region based.
I confirmed this by checking their nameservers, which are currently pointed at Dyn (ie, NS1.P24.DYNECT.NET).
The thing that amuses me about this is that Edgecast offers this same service at a far, far more reasonable price. I'm not sure why they didn't just use them.
""" At the highest level, a “balanced” CDN architecture is one that leverages two or more CDNs hosting identical content to increase (a) the overall number of physical Points of Presence for the network (PoPs) and (b) the proximity of those PoPs to the end users accessing them around the world """
I worked at a place that ONLY cared about the longest response time. Imagine! They ignored everything else!
What we are optimizing here for is end-user experience and minimizing the chance they have a higher-than-tolerable response time regardless of the variation they see.
Only then should you choose the statistic(s) you 'care' about.