Yes and no. Anything with a larger N is going to make for greater confidence in the rating. But that rating still depends on the number of positive and negative reviews. So if a product has a large number of ratings, but they are 50/50 positive and negative, then it's not going to rise to the top simply because it has the most reviews. The algorithm would presumably be very confident about placing it in the middle of the pack.
What I'm saying is that it will rise to the top because even though only 50% of the ratings are positive, it will have many many more positive ratings than the nearest next item. That 2nd item may have a much higher percentage of positive ratings, but a much lower number of total ratings.
require 'rubygems'
require 'statistics2'
def ci_lower_bound(pos, n, power)
if n == 0
return 0
end
z = Statistics2.pnormaldist(1-power/2)
phat = 1.0*pos/n
(phat + z*z/(2*n) - z * Math.sqrt((phat*(1-phat)+z*z/(4*n))/n))/(1+z*z/n)
end
puts ci_lower_bound(60, 100, 0.10) # => 0.517809505446319
puts ci_lower_bound(500, 1000, 0.10) # => 0.474027691168875
puts ci_lower_bound(50, 100, 0.10) # => 0.418847795168265
The top ranked result has 60 positives out of 100, and beats 500 out of 1000 - almost 10 times as many positive votes.Finally, I've got to say, 26 reviews (for the #2 item) doesn't seem insignificant, and the #2 item seems to have something like a 15% higher _average_ rating. Also, on another trending rating (http://www.rateitall.com/t-3239938-2010-ncaa-tournament-team...), the #2 has a significantly higher rating (17-18% it looks like), but only one less review: 4 vs. 3. I think if I were sorting either of these based on the average rating and the # of ratings, I would have done both differently.
Based on this, it seems that the Wilson score probably _over-emphasizes_ sample size, especially on things like ratings on high-traffic internet sites that may have orders of magnitude swings in the number of ratings.
It seems like the Amazon method of average works just fine, except for items with very few ratings possibly receiving disproportionately high ratings. Again, edge cases, which should probably just be penalized manually.