I was happy but surprised to find my startup http://taskarmy.com in the list, and having a closer look at the metrics taken in account, it seems that it is because of the StumbleUpon stats.
[lifx.co] => 97.942429714246
[kickfolio.com] => 89.598662998513
[shebusiness.com] => 84.413930181291
[netcomber.com] => 83.758506325309
[readershop.com.au] => 71.445369888231
[nameterrific.com] => 71.375673413085
[bugcrowd.com] => 71.328010040444
[shop2.com] => 71.228983543186
[serviceseeking.com.au] => 68.308687033257
[righttoknow.org.au] => 67.654116949018
[theiconic.com.au] => 67.28062258823
[kogan.com] => 67.173702738231
[retailmenot.com] => 65.963267045507
[startlocal.com.au] => 65.922497275844
[social-medicine.org] => 65.833298346458
[kaggle.com] => 65.616138632997
[manageflitter.com] => 65.551656760889
[freelancer.com] => 65.455719109403
[harris.com] => 65.150296026706
[freelanceswitch.com] => 64.742373047548
Doesn't feel like a substantial shift (good).But what is the logic behind the formula used for the fast_growth metric?
I see that the formula is: log(day14_sum_of_metrics/day1_sum_of_metrics) - 2*sqrt(1/day1_sum_of_metrics + 1/day_14_sum_of_metrics)
Is this a standard metric? Would be great if someone could share the theory behind the math..
Assuming that the first measurement is x likes, the second is y likes,
Then the estimate of the natural logarithm of the growth is given by
log(y / x) with an error estimate of sqrt(1 / x + 1 / y)
But since you are interested in the conservative estimate of the growth, you should use something like ~ 5% confidence interval. So I would recommend ranking your dataset using the folllowing function. log(y / x) - 2 * sqrt(1 / x + 1 / y)
For example:
growth from 1 to 10 will get the score of 0.2
growth from 100 to 400 will get the score of 1.16
growth from 10000 to 15000 will get the score of 0.38
One of the important properties of this estimator will be that the growth from say 10000 to 100000 will be ranked higher than the grown from 1000 to 10000, which in turn will be ranked higher than the grown from 100 to 1000 etc...
Coincidentally there is another post on HN's front page right now on Data Science resources: http://news.ycombinator.com/item?id=4930965
Matt's teacher's statement on the lack of knowledge on Linear Algebra was: ‘How can you make cheese if you don’t know where milk comes from!? Its plain, common ordinary horse sense!’
That hit home hard :(
On another note, Googling for: log(y / x) - 2 * sqrt(1 / x + 1 / y) throws up one heck of an awesome graph