How Not To Sort By Average Rating
evanmiller.org
evanmiller.org
For instance, here's another much simpler formula that avoids the problems described in the article: score = (pos+1)/(pos+neg+2). It's the posterior mean of Pr(random person likes the product), if your prior is uniform on [0,1]. You can adjust this to favour items with more ratings by, e.g., using a prior that "prefers" smaller values of that probability; one easy way to do that is to pretend that every item starts with a certain number of positive reviews and a certain number of negative reviews; you end up with a formula of the form (pos+A)/(pos+neg+B).
Is the article's formula better than this? Probably, but the author hasn't said why.
What I would love to see is a ranking algorithm that provides an elegant, intuitive interface to fuzzy ranking. If we can't be confident of the ordering, why pretend? It's a usability problem, not a statistics problem.
Your estimator is slightly more resilient than Amazon's and somewhat less conservative than the author's. It's definitely simpler, but who cares? Visitors never have to compute it after all.
(Yes, it would be extremely cool to make an elegant, intuitive interface to a fuzzy ranking. One reason why it would be extremely cool is that it's totally not obvious how it can be done, or even that it can be done.)
I wasn't suggesting that my estimator is better because it's simpler, nor that anyone should care; just pointing out that the author of the article made a huge leap from "here are two things that don't work" to "and here is the specific arbitrary math-heavy thing I think you should do instead".
http://code.reddit.com/browser/r2/r2/lib/db/sorts.py
s = score(ups, downs)
order = log(max(abs(s), 1), 10)
sign = 1 if s > 0 else -1 if s < 0 else 0
seconds = epoch_seconds(date) - 1134028003
return round(order + sign * seconds / 45000, 7)
Let me know if you want my attempt at a MySQL translation of that.ps. why is my account ( _ck_ ) so slow when logged in and my comments marked as "dead" ? Was I flagged as a spammer by mistake? Example: http://news.ycombinator.com/item?id=1219475
Looks like it. Looking at your posting history on _ck_, though, I don't see why you'd be auto-deaded. I suggest emailing pg about it.
Let's say you are ranking cellphones that a store sells and someone has a fond memory of a cellphone they used a few years ago so they vote it up. But the phone itself was posted years ago, compared to the newest phones that came out this year. You don't want their vote to have as much weight because of the product age.
But for actual product "scores" that people see that's a very bad idea. A book shouldn't have a 4/5 if everyone rated it a 5/5 just because it came out 50 years ago.
Minor: Distributions of ratings are not usually normal about the average possible rating -- they are usually skewed to the positive end, because people are pretty good at selecting what they read or view. The solution presented will tend to skew low-frequency ratings too far toward the average. It should be easy to calculate an expected average for your site, and use this instead of 1/2 for the formula given in the post.
Major: What you really need is a missing data point: the number of readers who haven't made any rating at all. It's pretty clear if you think of two examples:
Item 1 1,000,000 views 10 positive ratings 1 negative rating
Item 2 15 views 10 positive ratings 1 negative rating
Clearly, readers feel pretty blah about Item 1. You're much better off displaying Item 2, other things being equal, but the formula doesn't distinguish the two cases.
Would also be curious to see a comparison of different rating systems, for example i'd like to see what the distribution is on a site thumbs up/down rating site like digg vs a 5 star system.
Back to the original point -- if you used a formula with a default rating of 3 it would underestimate low-frequency movies. People don't view movies at random. They try to pick things they will like, which should be reflected in the statistical model.
> Indeed, it may be more useful in a "top rated" list to display those items with the highest number of positive ratings per page view, download, or purchase, rather than positive ratings per rating. Many people who find something mediocre will not bother to rate it at all; the act of viewing or purchasing something and declining to rate it contains useful information about that item's quality.
Yes and no. Anything with a larger N is going to make for greater confidence in the rating. But that rating still depends on the number of positive and negative reviews. So if a product has a large number of ratings, but they are 50/50 positive and negative, then it's not going to rise to the top simply because it has the most reviews. The algorithm would presumably be very confident about placing it in the middle of the pack.
What I'm saying is that it will rise to the top because even though only 50% of the ratings are positive, it will have many many more positive ratings than the nearest next item. That 2nd item may have a much higher percentage of positive ratings, but a much lower number of total ratings.
require 'rubygems'
require 'statistics2'
def ci_lower_bound(pos, n, power)
if n == 0
return 0
end
z = Statistics2.pnormaldist(1-power/2)
phat = 1.0*pos/n
(phat + z*z/(2*n) - z * Math.sqrt((phat*(1-phat)+z*z/(4*n))/n))/(1+z*z/n)
end
puts ci_lower_bound(60, 100, 0.10) # => 0.517809505446319
puts ci_lower_bound(500, 1000, 0.10) # => 0.474027691168875
puts ci_lower_bound(50, 100, 0.10) # => 0.418847795168265
The top ranked result has 60 positives out of 100, and beats 500 out of 1000 - almost 10 times as many positive votes.Finally, I've got to say, 26 reviews (for the #2 item) doesn't seem insignificant, and the #2 item seems to have something like a 15% higher _average_ rating. Also, on another trending rating (http://www.rateitall.com/t-3239938-2010-ncaa-tournament-team...), the #2 has a significantly higher rating (17-18% it looks like), but only one less review: 4 vs. 3. I think if I were sorting either of these based on the average rating and the # of ratings, I would have done both differently.
Based on this, it seems that the Wilson score probably _over-emphasizes_ sample size, especially on things like ratings on high-traffic internet sites that may have orders of magnitude swings in the number of ratings.
It seems like the Amazon method of average works just fine, except for items with very few ratings possibly receiving disproportionately high ratings. Again, edge cases, which should probably just be penalized manually.
It seems to me that the Bayesian average is superior, because it considers the value relative to other items.
The second example also isn't ordinal data—it's just nominal "liked" vs "didn't like."
Both examples are interesting, but don't necessarily get at the problem with ordinal data introduced by the article.
http://blog.reddit.com/2009/10/reddits-new-comment-sorting-s...
It's been a great resource for me.
EDIT: Fixed the link.
/me hangs head in shame.