A/B Testing Duration Data
evanmiller.org
evanmiller.org
That on top of the unwarranted distribution assumption is very dangerous. How dangerous? If the true distribution has heavier tails than assumed, then the true variance can be many times what is assumed. If the true variance is, say, 4x what is assumed then an average standard deviation variation will show up as passing a 95% confidence interval in a random direction. Getting to 99% confidence by accident becomes quite easy, with bad results.
However the difference between the t-test and the normal to which it is an approximation is that the t-test takes into account corrections the rate of convergence to the estimated normal. The details of that convergence is dependent upon the distribution of individual samples, and the t-test is predicated on the assumption that each and every sample is from a normal distribution.
Terrible assumption. It's probably something closer to this:
P(Leaves site without reading / sub 2 sec): 0.2
P(Leaves site after skimming / sub 10 sec): 0.6
P(Leaves site after reading 1 page / 5 minutes): 0.1
P(Leaves site after reading a bunch of pages / 30 mins): 0.1
If you're going to make decisions off of something like duration, then you should code up a basic bit of javascript/database that can store real values.http://snowplowanalytics.com/analytics/catalog-analytics/mea...
(You can't measure visit durations including bounces accurately without a JavaScript tracker that pings back on-page DOM activity.)
I'd stick with assuming duration time is normally distributed. This leads to another class of adhociness: http://www.statit.com/support/quality_practice_tips/estimati...
If anyone is interested in trying it out, my email is in my profile. Always looking for feedback.
[1] Actually I'm waiting to build it. I quit my job a couple of days ago and will begin work after my notice period is up.
Run something like this through 10,000 simulations and you'll start to see how the theory really gets applied.
One very long (say 2 days) visit will skew the average. And you can't infer the standard deviation from the average.
For an easy alternative, you could calculate the median time on site ex ante -- let's say it's 1 minute. Then for your ab test, see what percent leave within the first minute. Use that percent and determine if difference is statistically significant.