How Not To Sort By Average Rating
evanmiller.org
evanmiller.org
Assume that ratings are being generated by a stable stochastic process where the underlying distribution is multinomial (ignoring the ordinal character of ratings, for the time being) and use a dirichlet conjugate prior. This gives you a posterior distribution over new ratings for an item. The benefit of a posterior here is that it lets you rank items by thinking in terms of the probability that the viewer would rank one item higher than another at random. By adjusting the magnitude of the alpha parameter to the dirichlet prior, you adjust your sensitivity to small numbers of observations. A small initial alpha will lead to rapid changes in the posterior upon observing ratings, whereas a large alpha requires a significant body of evidence.
The best part of the multinomial model with conjugate dirichlet prior is that the math is REALLY simple. The normalizing constant for the dirichlet distribution looks scary when stated in terms of the gamma function, but given this is the discrete case, just pretend everywhere you see the gamma(x), it is replaced with (x - 1)! and you will be ok.
Let me know if you would like to learn more, I would be happy to help.
Is there any other sort of case?
I emailed Miller a while ago to see what he thought of this reply, and he thought it also seemed like a reasonable approach. But, in his view, the criticisms of his method within their framework include things that in practice he sees as features. In particular, they view the bias caused by using the lower bound as a bug, but he prefers rankings to be be "risk-averse" in recommending, avoiding false positives more than false negatives. Of course, that biased preference could also be encoded explicitly in a more complex Bayesian setup, which would also be a bit more principled, since you could directly choose the degree of bias, instead of indirectly choosing it via your choice of confidence level on the Wilson score interval.
The strength of priors here is that it is very easy to take intuitions and encode them statistically, in an understandable way. Taking the lower bound of a test statistic doesn't admit much in the way of intuition.
That makes for a pretty significant difference in practice, and I'm not sure which is preferable.
In all of these models, the giant variable that is completely ignored is the actual choice to rate something at all, versus skipping over it and reading the next one. That's a very significant decision that the user makes. The behavior of each of these systems w.r.t that effect will be the dominant thing differentiating them.
Here are slides from that paper's presentation, for a quicker overview: http://www.dcs.bbk.ac.uk/~dell/publications/dellzhang_ictir2...
Or you could do a semi-frequentist thing and simplify your math by using MAP estimates to rank. Basically instead of score = #pos/(#pos + #neg), it becomes score = (#pos+x)/(#pos+x + #neg+y), where you choose x and y to suit your needs. You could choose x/y in proportion to the average number of up/down votes on your site or you could even choose x/y in proportion to the average number of up/down votes of the author of the post. That would rank posts of trolls lower than posts of good users. By varying x and y you can tweak the strength of this effect. You can interpret this as giving each item by default x upvotes and y downvotes.
This certainly works much better than the formula in the article. For example if a post has 1 upvote and 2 downvotes, his formula will say that should be ranked lower than a post with 1000 upvotes and 2000 downvotes (because he's using the lower bound of the confidence interval). Obviously that's bad because while the first post could be a good one, we know for certain that the second one isn't. In general his method will rank posts with a low number of votes very low, even compared to posts with a high number of downvotes.
The benefit of the Bayesian treatment here that I want to drill down on is how natural it is to adjust the prior to capture your beliefs about how items should be perceived in the presence of incomplete information. The frequentist approach is fine, but it does not provide such a pleasant, intuitive knob to tune.
But I agree that the Bayesian approach is conceptually much cleaner. IMO the frequentist approach is just computational corner cutting for when the math in the Bayesian approach gets too involved, which is sometimes useful. What's nice about the Bayesian approach is that you state your assumptions and then it's just turning the math machinery. In contrast, in the frequentist approach the assumptions are interwoven and hidden in arbitrary choices in how the math is done (And then they claim that Bayesians are subjective! It's just that Bayesians admit that they are subjective. Frequentists try to hide the fact that they are more subjective in the math). The not so nice thing is that turning the math machinery is not always so easy and does not always produce fast algorithms. That's where maximum likelihood and friends come in, but I'd view them as an approximation to Bayesian methods.
I love Bayesian statistics from a conceptual point of view, but the ease with which one ventures into the land of analytic intractability kind of puts me off more complex models. MCMC is such a clumsy tool (in addition to taking forever); variational methods look interesting to me but I don't really feel they are quite there yet.
In this context, the pessimistic prior is more of a hack than the confidence interval, because your mental model isn't that most comments suck. Your mental model is that most of them are ok. The pessimistic prior is just there to say, if 1000 people confirm that a comment is ok, we should show that over a comment for which we don't know whether it's ok or not yet, and this reasoning is much better modeled by a lower bound.
Do you have any recommendations for a good introductory treatment of Bayesian statistics?
http://www.inference.phy.cam.ac.uk/mackay/itila/book.html
Videolectures has some very good videos as well. Zoubin Gharamani has a pretty solid lecture on Bayesian learning at http://videolectures.net/mlss05us_ghahramani_bl/ (he's a great researcher but not the most engaging speaker). Try Christopher Bishop's lecture at http://videolectures.net/mlss09uk_bishop_ibi/ as well, it might be slightly more palatable.
His book starts from first principles -- simple ideas about probabilities -- and it builds a foundation for understanding Bayesian methods. And he explains, basically, why the p-values and t-test stuff is a bunch of crap, which was an immense relief.
Thanks for the resource, this looks fantastic.
Why I love it: It's precise. It's elegant. It's rigorous. It's based upon solid, proven science & theory. It's a perfect application for a computer. And most of all, it does what's intended: it works.
Why I hate it: What human can understand it?
I used to implement the first manufacturing and distribution systems that used thinking like this. They figured, "We finally have the horsepower to apply complex logic to everyday problems." Things like safety stock, economic order quantities, reorder points, make/buy decisions, etc.
But the designers of these systems overlooked one critical issue: these systems included humans. And as soon as humans saw that decision making formulas were too complex to understand, they relieved themselves of responsibility for those decisions. "Why didn't we place an order?" "Because the computer decided not to and I have no idea why."
I suppose the optimal solution is somewhere in between: a formula sophisticated enough to solve 95% of the problem but simple enough for any human's reptile brain to "get it". This isn't it.
I talk about this with some of my med school friends that are interested in/want to create/hate/fear automated diagnosis. At one level, having appropriate statistics to make use of a wealth of prior experience, worldwide prevalence, epidemiological data, &c is basically a requirement for the future of proper healthcare.
It's also an obvious terrible mistake to ever convince people that these statistics are anything besides a decision making tool.
I think the right usage of statistics is to enlighten and confuse simultaneously, not to "answer". They should provide analytical depth to a decision, never an escape route.
For this reason, I think proper statistical application is an intersection between mathematics, computer science, and UI design.
You're right, but I would go further. I would argue that when statistics are presented as "the answer," in many cases they are being abused. A good scientist recognizes the limits of his dataset. Like how a computer program can only do what it is written to do, statistics can only tell you what is in the data, they can't tell you what is _not_.
There's nothing wrong with the math. It just requires better educators to explain it, with analogies and metaphor. Why would you compromise the data/outcome in a hope to simplify the problem?
Further: wouldn't you look to hire (and train) the best people who understand the domain they are in, thereby being able to judge whether the outcome of an equation is valid or not?
Hint: insurance/loan underwriters regularly end up in situations where the human component of a transaction may look different than the data, and react accordingly...
So... anyone want to take a stab at explaining that equation to those of us who don't really get it?
[1] http://blog.reddit.com/2009/10/reddits-new-comment-sorting-s...
Having figured out the midpoint c, Wilson has to actually figure out the lower bound x and upper bound y. Now he draws a distribution centered at c....ok so at this point you would need to know what is a distribution, and why you would need one, whether that distribution has a skew & whether its homoskedastic & so on...which is stats 101, so I won't go there. But if you've gotten this far, you should be able to atleast see the intuition behind Wilson's procedure.
- You have a normal distribution (bell curve) of data points, in this case quality scores.
- You wish to sort these points based on their respective vote totals. Any given data point has pos positive votes out of n total votes for that item.
- You have a confidence interval, e.g. 95%. This confidence is expressed in terms of the bell curve, so a 95% confidence is within 1.96 standard deviations of the mean [2].
- You have a Ruby function accepting the aforementioned n, pos, and confidence variables and returning a decimal value representing the normalized confidence_interval_lower_bound, that is the quality score that our input data point has a 95% chance of meeting or exceeding.
- Given a set of data points, evaluate the ci_lower_bound for each, and then sort them accordingly. The results will give you a best-guess sorting that accounts for the fact that some data points will have more votes cast for/against them than others.
[1] http://www.med.mcgill.ca/epidemiology/hanley/tmp/Proportion/...
Now say your users vote, and they upvote with that probability p. The number of upvotes k you get out of n votes will follow a binomial distribution B(k;n,p). The binomial distribution has mean np and stdev sqrt(n p(1-p)), and is very close to gaussian in shape. Since the stdev is a rough measure of the 'width' of the distribution, common way to describe the error is (mean +/- lambda*stdev), where you can tune lambda to your desire. If you increase lambda you get a wider confidence interval, and therefore more certainty that a measurement will be within that confidence interval.
Now, say you measure k upvotes out of n votes. You can divide by n to get p0, your estimated rating based on those votes.
An easy estimate for the error of this measurement is to assume that p0 is approximately correct and equal to p. Then the expected number of upvotes would be n p0 with stdev(n p0(1-p0)). Divide this by n to get the fraction of upvotes, to give a final estimate for p of p0 +- lambda sqrt(p0(1-p0)/n)
Now, your estimate (p0) of p is not quite right, and therefore your estimate of the error (which depends on p0) is not quite right either. The wilson score attempts to correct for that. We don't know p, but imagine if we did, we would expect any measurement p0 to be in the range p +- lambda sqrt(p(1-p)/n). That is, we expect
abs(p - p0) < lambda sqrt(p (1-p)/n)
If you solve this equation for p in terms of p0, you get a formula given in the article, ie, the confidence limits for p given p0.I never said there was. In fact, I praised it as an elegant solution.
It just requires better educators to explain it, with analogies and metaphor.
You're right. In theory. In practice, no one does this, mainly because they can't afford it. You're implementing technology that costs $200,000 to save $100,000.
Why would you compromise the data/outcome in a hope to simplify the problem?
Actually, simplying the problem actually reduces the compromise when humans are involved. So, no.
...wouldn't you look to hire (and train) the best people who understand the domain they are in, thereby being able to judge whether the outcome of an equation is valid or not?
Only if it made economic sense to do so. I am not going to replace 800 workers earning $12/hour with "the best people who understand the domain" because the computer suddenly has formulas that people don't understand. Experience has shown repeatedly that workers, at any level, simply stop caring when they feel powerless by "solutions" like OP's Wilson formula. That abducation of responsibility almost always far outweighs any incremental benefit that a more sophisticated but incomprehensible formula introduces.
Good points, nice discussion, but please, for the sake of this community, don't reply with words like "what?", "Further:", or "Hint:". You made your point without the snarkiness.
Yes, assembly line workers need simple steps to do a job. But that of course isn't who this is referring to. In your example managing inventory isn't an assembly line, checklist job.
You're not replacing 800 workers with the best. In your example, you wouldn't have 800 people managing inventory. You replace your inventory manager with someone who understands the math (if you want to manage your inventory that way)
I think that's the point edw519 was trying to make.
Add in a UX that makes it obvious that there's a choice to be made and allows the user to easily evaluate the results of the naive solution alongside the optimized solution, and end users who are themselves evaluated on overall performance metrics will tend to choose the more effective solution whether they understand it or not.
The addition of metrics to allow for side by side evaluation of the simple strategy versus the optimized strategy is key - someone choosing the unintuitive yet superior strategy will likely have to explain themselves to management at one time or another.
Aan interesting parallel is NFL onfield playcalling decisions. There are plenty of instances where coaches make choices that are statistically unsound but they reinforce the image of a tough, responsible coach. Conversely there are coaches who play by the odds and get lambasted when their decision doesn't pan out.
And somehow the solution isn't to have the computer output a detailed description explaining why it arrived at the decision that it did?
When you're working with a system managing orders, where there's a direct human interaction with the algorithm (like the top-level poster is), simplicity is somewhat important. When you're working with a system that's designed to look like magic to the end user, making the system more magical is often better, because it makes it less intuitive to game.
Besides, the math made no effort to write the sentences. ;)
Webmasters * constantly* complain about the impenetrability of the 'algorithm' and how it constantly changes for secret reasons.
So you're right in a way---the demand side doesn't care how you manage to curate supply, but supply is intimately aware of your curation and worries about being unfairly demerited.
I think as comments get higher to the top, people start voting them down more, which might not be the cases for comments which are rising (theory: people reading further down in the comment section are possibly more thoughtful and more likely to upvote, while at the top the ADHD crowd might be more likely to knee jerk downvote?).
In any case, i've recently changed my default sort to "top" and feel it's an improvement.
If you think this is too difficult then maybe you shouldn't try to design scoring/voting systems.
This just means you need a better notification system, not a dumber model. I am fairly sure Netflix and Pandora use complicated recommendation algorithms - we don't complain because the complexity is abstracted away behind labels like "dark cerebral thrillers" or "songs with rhythmic gyrations like Pretty Lights".
The formulas aren't "too complex to understand", the people are being too lazy to take the time and understand them. No reason to dumb it down and ruin it for the rest of us.
Lets start with Wilson's midpoint, since that's just high school math.
def mid(upvotes:Int, downvotes:Int) = {
val total = upvotes+downvotes+0.0
val up = upvotes/total
val half = 0.5
val a = total/(4+total)
val b = 4/(4+total)
a * up + b * half
}
So there are two weights a and b. Using these weights, the midpoint is a weighted average of half and the proportion of upvotes.It should be very clear that if the total becomes large, a goes to 1 and b to zero. At that point you end up using the proportion of upvotes, just like Amazon.
Now, lets bring in the entire confidence interval.
def wilson(upvotes:Int, downvotes:Int) = {
val z = 1.96
val n = upvotes+downvotes+0.0d
val phat = upvotes/n
val lower = (phat + z*z/(2*n) - z * sqrt((phat*(1-phat)+z*z/(4*n))/n))/(1+z*z/n)
val upper = (phat + z*z/(2*n) + z * sqrt((phat*(1-phat)+z*z/(4*n))/n))/(1+z*z/n)
(lower,upper)
}
Using the author's data points, the results look like this: val itemVotes = List((600,400),(5500,4500),(2,0), (100,1))
itemVotes.foreach( x => {
val s = mid(x._1,x._2)
val w = wilson(x._1,x._2)
printf("Up:%5d\tDown:%5d\tMid:%.3f\tW:[%.3f,%.3f]\n",
x._1, x._2,s,w._1,w._2)
})
scala> Up: 600 Down: 400 Mid:0.600 W:[0.569,0.630]
Up: 5500 Down: 4500 Mid:0.550 W:[0.540,0.560]
Up: 2 Down: 0 Mid:0.667 W:[0.342,1.000]
Up: 100 Down: 1 Mid:0.971 W:[0.946,0.998]
Basically, we prefer the lower bound of the confidence interval instead of the midpoint of that same interval.Case in point, the article mentions Amazon. Amazon could easily hire the best and brightest mathematicians and developers to make the implementation completely rigorous and provably correct. However, they're still going to field questions from merchants saying "My product has a 5-star average, but this competitor's product only has a 4.5-star average. Why the hell is my product shown further down in the list?"
Similarly, you may work in the acquisitions department of BigCorp and you need to purchase some equipment. When you prepare your report for Big Boss of your recommendations, how do you explain that you are recommending something which has a 4.5-star average rating instead of the 5-star average rated one?
The ratings may be 100% mathematically sound, but the users of the rating system are not the same as the producers of the rating system. So it's largely irrelevant if the producers made it perfect if the users don't understand what "perfect" means.
for amazon, is it better to show the best product first or to help sell a few of the lacking-rating ones to generate more ratings and work out the uncertainty with real world data instead of crazy math?
You don't need to let them rate with stars and also show it as stars.
'Because you have fewer votes than your competitor'
I don't think it's that hard to understand.
I disagree that people won't be able to understand it - I propose they can't be bothered to understand it. Regular people do not give a shit about algorithms - either the computer/software works or it doesn't. Tech geeks regularly forget this simple fact.
Most people have other things they care about and all they want is to press a button and have an answer come out.
If its instantly clear how the answer is generated, they think about it for a very short while, agree and move on with their lives.
If it's not instantly clear how the answer is generated, they assume the computer is right and move on with their lives.
This may or may not be a problem. Remember, human intuition evolved to be able to give a fast answer to any question in the face of insufficient data, not an optimal answer to specific questions with sufficient data. In many domains, once you have appropriate data and the means to statistically analyze it, an algorithm using such analysis can reliably outperform human experts - and in some cases, the problem was the political one of how to get the human experts to accept this, take their hands off the wheel and stop second-guessing the algorithm.
In other words, if you are dealing with such a domain, you may be better off to count your blessings, let the computer do its job and save the humans for tasks that can't be handled by a formula.
http://www.dcs.bbk.ac.uk/~dell/publications/dellzhang_ictir2...
Of course, I think the authors miss the point of the algorithm, since I basically wanted a system that is one-sided (i.e. false negatives are OK but false positives are bad).
Also, if you deal with more than two outcomes you might be interested in multinomial confidence intervals, described here:
http://www.math.wsu.edu/faculty/genz/papers/mvnsing/node8.ht...
The application to 5-star systems is not straightforward, since it's not clear to me how stars relate to each other. Is it a linear scale? Are they discrete buckets? Or maybe we want to use Tukey's froots and flogs? I'm not sure.
By the way, I'm coming out with a stats app for Mac soon that implements this algorithm and much more. Drop me your email address if interested:
5 star rating systems are obnoxious. From a mathematical perspective, if you treat them in an ordinal fashion they are poorly behaved, and if you treat them categorically, you lose the relationship between stars. There seems to be some popular movement towards binary rating systems, and I think that is great. Not only do people tend towards binary rating behavior in the real world (only rating a movie they thought was very good or very bad) but they admit a much cleaner mathematical treatment.
Helping out a friend with a statistics test, I recently read up about the Wilcoxon Signed Rank Test[1]. Now this one is intended to get a p-value for experiments with "before" and "after" measurements, but what I got the idea it's trying to do, is to use the ranks of a not-very-normal behaving random variable, turn that into a summation of lots of things, so due to the central limit theorem you can treat it as a normal distribution again.
Though thinking about it, in this case it's the rank we're after, so maybe it's not useful at all. But it gives an interesting idea about the tricks you can pull if your input data isn't quite the sort of type that you can analyse very well.
The main thorniness of 5 stars is that you have to answer the question of what the difference in star ratings actually mean. Is going from one star to two stars the same as four to five? Probably not, based on the way users rate, which means an algorithm like the arithmetic mean that treats them the same is wrong.
Personally, I think it's very reasonable to treat rating stars the way we should treat grades: as ordinal data, where we know that a higher rating is better but assume nothing beyond that. The difference between an F and a D is not the same as an A and a B, and the same is likely true of 1-2 stars and 4-5 stars.
I have made an attempt at applying this idea to the Ubuntu App Store's rating algorithm. I'm very much interested in comments. https://bugs.launchpad.net/ubuntu/+source/software-center/+b...
Typically, you do not know that, especially if you are comparing across reviewers. Some will only score zero and 5 stars, other will have 10% two stars, 80% three stars, and 10% four stars, yet others will have 10% three stars, 80% four stars, and 10% five stars.
If you have sufficient data (rare), it may be possible to (somewhat) correct for that. IIRC, this was something that helped in winning the Netflix challenge.
For example, http://www.netflixprize.com/assets/GrandPrize2009_BPC_Pragma... models the assignment of stars as two parts:
- modeling the user's appreciation of the movie - modeling how the user translates his appreciation to a star rating
"where we know that a higher rating is better but assume nothing beyond that"
Problem with that is that you throw away information with that assumption. You do know that 2 star scores are very unlikely to be about very good items.
Would it be reasonable for 5 stars to normalise the data? Should star ratings be on some distribution, for instance?
This is a little flakey for ordinal numbers, and the usual treatment is to use a learning algorithm to find a mapping from real numbers to ordinal values, either explicitly (if you need a "score") or implicitly. Support vector machines, radial basis functions and neural networks are typically used.
The assumption is that there's some constant p underlying probability that a random person will rate a given thing positively. If we observe, for instance, 4 positive and 5 negative reviews or votes, there's a probability distribution (known as a Beta distribution) which tells us what the possible values of p are given the votes we observe: p^4 (1-p)^5. graph: https://www.google.com/search?q=x%5E4+(1-x)%5E5%20from%200%2...
Now if we observe 40 and 50, respectively, the curve looks like this: https://www.google.com/search?q=exp(20+%2B+40+log(x)+%2B+50+...
(I had to do it in the log domain because Google's grapher underflows otherwise -- the 20 is just to make the numbers big enough to graph. The more correct thing involves gamma functions and that just gets in the way right now)
The more you observe, the more sharply peaked the likelihood function is. The funky equation in the article is an approximation to the confidence interval of that graph -- 95% of the probability mass is said to be within those bounds.
It's not a great approximation, for one because the graph is skewed (try it with 10/50) and it assumes that the mean is exactly in the middle of the confidence interval. The correct computation involves the inversion of a messy integral called the incomplete beta function. Scipy has a package which includes betaincinv which solves this more exactly:
>>> import scipy.special
>>> scipy.special.betaincinv(5,6, [0.025, 0.975])
array([ 0.18708603, 0.73762192])
would be the 95% confidence interval for 4 positive and 5 negative votes;
>>> scipy.special.betaincinv(41,51, [0.025, 0.975])
array([ 0.34599562, 0.54754792])
for 40 and 50, respectively.
[edit: apologies, I had to run and get ready for work -- I didn't really have time to make this very comprehensible; but i just now fixed a bug in my confidence interval stuff above]
And to make the description slightly more accurate, at the expense of more complexity: "What number are we 80% certain the approving percentage will exceed?"
Anyway, I personally think 95% confidence intervals are a crutch. The correct Bayesian approach is to consider two items, each with their own up and down votes, and integrate over all possible values for p1 and p2 (being the underlying probabilities of upvotes for item 1 and 2, respectively) over the observed data, and compute the likelihood of superiority of p1 over p2.
How to turn that into an actual ranking function? No idea. I doubt it would work, but you could compute against a benchmark distribution (i.e. the uniform 0-1 distribution).
If you do that, it probably turns out that your ranking function is the mean of the Beta distribution, which is simple: (U+1)/(U+D+2) where U and D are the upvote/downvote counts [note: we started with the prior assumption that p could be anywhere between 0 and 1, uniformly]. Basically, the counts shrink towards 1/2 by 1. This is a hell of a lot less complicated, and it achieves the goal of ranking different items by votes pretty well with more votes being better.
The beauty of the second algorithm for rating products is that it is straightforward. Having never seen it before I can deduce that 5 stars come before 4 stars and more reviews come before fewer. If I want to skip ahead to the 4 stars I know what to do. I can internalize the sorting algorithm easily. And as a user, understanding the order items are presented to me is important.
If Amazon were to use the last algorithm and present items in that order (assuming we accounted for the 5 star vs positive/negative issue), it would like a random order to most users and would be frustrating.
So I guess what I am saying is that this algorithm is very clever, but in some cases, it may be too clever. Sometimes you just want to keep it Simple Stupid.
80-100% = * * * * *
60%-79% = * * * *
etc..
Currently, when we see an item's star value, we think of it as an indicator of the quality of the item. But if it's just the average of the star values of every review, the author would argue that we're not going to get an accurate indicator of quality. The author argues that whether the quality indicator of an item is expressed in stars or percentages, that value should be determined by the third algorithm, not the second, and that the order the items are shown in should be the result of sorting those quality indicators.
Instead you could use a weighted baysian rating:
br = ( (avg_num_votes * avg_rating) + (this_num_votes * this_rating) ) / (avg_num_votes + this_num_votes)
The point is, I understand the sorting order and can manipulate them if I am not satisfied with what is presented to me. Having a very esoteric algorithm is a risk. Maybe you'll present just what the user really wanted. But if you get it wrong they will be lost to do anything about it. I tend to dislike systems that leave users helpless when something goes wrong.
One simple fix would be to avoid calculating an average until a minimum number of ratings have been given. But I do think the statistical way is lovely. If I were Amazon I'd give it some kind of snappy trademarked name and push it as a feature.
(pn) / (n+q)
This is simpler and gives similar results:
http://www.wolframalpha.com/input/?i=plot+%28p+%2B+z^2%2F2n+...
http://www.wolframalpha.com/input/?i=plot+%280.8+*+n%29+%2F+...
Just sort the objects in the order determined by this formula and only show the ratings given by users in the interface.
The people reporting the problems were the users, and yes, this was bias, as they expected averages. That was exactly my point — while the statistics behind this method are sound, it is not what people expect. Building systems that don't do what people expect is difficult.
But... so what? Amazon and Urban Dictionary are hardly failing in the market due to their "incorrect" score sorting. The whole problem is a heuristic, it's not amenable to rigorous treatment no matter how many giant equations you club your audience with.
(I agree with other commenters that it is complicated and lacks common sense to average users, but I feel like I have a general understanding of the concept thanks to the above link)
However, if you install the stats package for PHP you can use stats_cdf_normal() as well.
http://www.derivante.com/2009/09/01/php-content-rating-confi...
I'm not a statistician so I can't speak to the correctness of this implementation.
I've found, in practice, a Bayesian weighted average is easy to implement and pretty effective. It's also a good candidate for "stream" processing (i.e., calculating in a single pass)
Or maybe more notably, https://github.com/reddit/reddit/blob/master/r2/r2/lib/db/_s...
- Considering the standard deviation of ratings. On a 5 point scale, an item that rates 3 because ratings are split between 1 and 5 votes, differs from one that gets mostly 3 votes. The latter is a middlin' fit for anyone, the former has an enthusiastic but niche audience. If you're looking at sales, the former can be a valuable product if properly marketed.
- An item that gathers few votes regardless of favorability ratings can exhibit multiple problems. One is that it isn't well marketed / publicised, or known. Another (particularly on content sites) is that there's very likely a sampling bias (mutual admiration society / negging attack / vote stuffing). I've tended to favor systems which take into account the total volume of voting, generally on a ln(n) basis, though not out of any particular statistical rigor. As an implementation, you'd start with a 5 point Likert score, then multiply by, say, ln(n+1) (avoiding a zero multiplier on a single vote).
- The pattern of ratings over time and space (IP or geographical) may reveal both opportunities for marketing and/or issues with your ratings system. Since any effective quality proxy will be abused, you've got to be sensitive to the latter.
The Wilson score is an improvement over multiple other methods. It still does assume a relatively unbiased estimator and rating behavior. My feeling and experience is that excess reliance on any one metric is likely to cause problems -- reality is multidimensional, metrics for assessing reality should be as well.
There's also the question of whether or not you want to make specific recommendations for an individual, or general recommendations for a population. In the former case, correlating other rankings or behavior may give a better fit (and the Wilson score may still be useful).
Though for a suitably specific goal (marketing, suitability, revenue potential) a single encompassing metric may work.
http://news.ycombinator.com/item?id=1218951 <- 31 comments
http://news.ycombinator.com/item?id=478632 <- 56 comments
Further, I hope JoshTriplett (http://news.ycombinator.com/user?id=JoshTriplett) isn't too disappointed that when he submitted this exact same item 2 days ago it got one upvote and no discussion. In submitting to HN, as with comedy, timing is everything. http://news.ycombinator.com/item?id=3784912
"USA: The only country keeping penguins from coquering [sic] the Earth"
This has been the number one complaint I have against Amazon for the past 10 years. And they haven't done a think about it?
This algorithm needs to be paired with another algorithm that weights each plus one according to each user's ability to plus one something that gets a lot of plus ones.
I couldn't hope to do the math for something like that, but I'd sure like to talk to someone that could.
http://www.amazon.com/Stand-Mixers-Small-Appliances-Kitchen/...
(TotalScore - 1) / MaxPossibleScore
Such that (using the Amazon examples from the article):
((2 * 5) - 1) / 10 = 9/10 = 90%
((100 * 5) + (1 * 1) - 1) / 505 = 500/505 = 99%
As a civil engineer working for a local city he might be involved in a 10year process of approvals to add a freeway on ramp. Where most of his job would be checking that an army of subcontractors were all doing things to code - not that they were doing things well, just to the written requirements
In Afghanistan if they want a road or a barrier he basically finds somebody lower rank points at a bulldozer and tells them to do it.
An interesting point he made was building a simple village clinic with a clean water supply that would save lives for a few days work and a few $1000. At home he would be involved in a multi $100M, 20year project for a new hospital where most of the money would go into pretty decoration and parking structures and would probably end up costing lives compared to the existing old hospital that was working perfectly well.
From: the boss
To: dev3
Subject: URGENT - front page showcase selection broken!!
Body: Hey bro, I was looking into it, and our ratings average equation is totally busted and products with just a few ratings are hogging space from proven winners when it's just a sample bias. This is costing us money and needs to be fixed NOW.
I'd like this up before our morning meeting so I can boast about it and you'll get credit too, as this should massively increase our conversions right away by putting BETTER products right on the front page.
this should get you started: http://evanmiller.org/rating-equation.png
I'm sure you'll figure it out. If you could do an A/B test for bragging rights too that would MASSIVELY rock. Thanks!!!
Rock on,
Boss