Brewing a Better Rating System
blog.steepster.com
blog.steepster.com
Background: Steepster is a community site for tea drinkers to share their tasting notes, get recommendations, and discover new teas.
Feedback appreciated!
1. is your slider from the jQuery UI or other js framework?
2. in regards to combating the 4.3 dilemma, have you found the average ratings on steepster to be lower? maybe its too early to tell, but I'd love to see some sort of curve on your ratings distribution in a future post ...
thanks
1. Yep, slider implementation is jQuery UI.
2. It is too early to tell, but we're definitely planning to share a follow up. As mentioned in the post, we had a simple thumbs up/down for ratings and were seeing a greater than 90% positive average, so we were definitely experiencing that bias. Just today, albeit with a much too small sample size, we're starting to see a more diverse mix of averages. We still expect to have that positive skew but because we're now operating with a 100 point scale in the UI, we hope the granularity will help users distinguish subtler differences in rating.
But, this is a good point, and I think an important one to consider when evaluating the mechanic that works best for your community/site.
Have you received enough ratings with this new interface to know if my assumption of ratings being clustered is accurate?
You've given us good food for thought, thanks. My general feeling now is that I think it's important to leave the negative portion of the slider intact (however less it's used) to maintain a solid mental model. Might be something to test down the road though.
Have you tested the theory that response bias is skewing the ratings upward?
In other words, given that not every tea drinker is going to bother logging on to your site, finding each tea they've tasted and leaving a rating, it seems plausible to me that people who've had a good tea-drinking experience are more likely to make that effort than those who've had an unremarkable tea.
The figures that you quote seem to support this theory. With a yes/no rating system you had 90% yes votes, or an average vote (assuming one yes vote cancels out one no) of 0.8, skewed 80% up from the unweighted mean you'd expect. With a 5-star rating you expect an average rating of 4.3, which is 65% up from the unweighted mean ((4.3 - 3.0) / (5.0 - 3.0)). So adding granularity is decreasing the skew, but not very rapidly. I'd be very interested to know what your averages are like now, with your new system.
Granted, some of the skew is due to what teas are available to be rated: presumably people are less likely to enter awful teas into your system in the first place. I realise the point of your article is about redesigning the rating to combat that skew, rather than necessarily about finding the one true rating. But if you're concerned about bias, it seems worth at least investigating all possible sources of bias.
There's an obvious asymmetry in this hypothesis - why would the response bias be in favour of strong positive experiences, rather than just strong experiences in general? Even if drinkers of mediocre tea can't be bothered to vote, why wouldn't people who've had terrible tea be just as likely to vote as those who've had great tea? I can think of two explanations. One is an innate sense that a good experience is worth more effort than a bad one - so after a bad pot of tea, there's less of an impulse to run off and tell everyone how bad it was, more to just write it off and go do something else. Another is that people motivated to tell people about bad experiences might want to do so in words, to explain what was so bad about it.
I realise this is all hand-waving - I don't have hard evidence to back up this theory. I do think it would be an interesting theory to test, for those with sufficient levels of usage of a rating system to do so.
Some anecdotal evidence comes from my use of other UGC review sites, particularly restaurant reviews. These sites usually have a disproportionately large number of negative reviews. Reading reviews for nearly any restaurant leads you to conclude that restaurant has terrible service - apparently because those who've had bad experiences are always keen to vent about them. Yet nearly all restaurants, even those with tens of vocal unhappy customers, have above-average ratings.
I'm imaging a UI which asks you to pick a favorite among the item you are viewing and one similar item. You could stop there, or ask repeatedly with new comparison items until the viewed item's position on the absolute scale is unambiguous. The user could provide some rating data with just a single binary decision, but some ajax-y fade out/in of another pair could enable further ratings if they desired.
Steepster should use a Bayesian average (http://en.wikipedia.org/wiki/Bayesian_average) so that the uncertainty of a small number of ratings is reflected in the sorting.
Take movies, for example. They are usually rated on a four-star scale. And yet, a three-star movie is a clear success. Few movies can realistically aspire to more than three stars. Even many four-star movies are really just trying desperately to avoid two-star land. Francis Ford Coppola was sure he was going to be fired any day from The Godfather. The production crew and actors on Star Wars thought it was practically a joke. Please, God, let Star Wars not be a B movie, they must have been thinking.
When you say ★★★ out of ★★★★, you make it look like it wasn't good enough: 75%. Movies really should be rated on a three-star scale: ★★★ out of ★★★; ★★★ = A = 100%. Anything else is gravy.
So, rate tea on a three-star scale. Three stars means "excellent tea, no clear way to make it better". ★★★½ means "Whoa, there is something better than ★★★!" ★★★★ means "This is The Godfather of tea! This tea makes me an offer I can't refuse."
The notches kind of do this but theres always the risk of someone rating there first tea 80, then deciding subsequent teas after are better so they need to be rated higher, when the first one should have been more around 60.
I didn't vote down because I disagree, but because you haven't made much of a case. I also downvote the 'wow, cool' comments as unhelpful, and upvote the comments that seem like they will lead to useful discussion. Without intending offense, I didn't think your comment was pitched at the right level for this audience.
Personally, I think you are on the right track, although I think 5 stars allowing halves is even better. Interestingly, Netflix (experts in this field) started out with allowing half-stars and then got rid of them, making me worry that they know better than I.
As for the 7 point system being "best" (for general purpose rating), I remember it's because 5 star does not give enough granularity, while 10 points is too much. Maybe that's why Netflix took that out. Now why not 6, 8, or 9? I forgot, again, maybe there was research being done.
As for my "idea" of allow a 0 on a 5 point system; it makes it a 6 point system while retaining a 5 point UI that everyone is accustom to. What is wrong with that? Again, just asking for discussion, not trying to say it IS the best.
Now back to the topic of down vote because I don't have enough backing. If I need backing in order to comment, I wouldn't even be able to comment any of this. Is this what you think is the way hn should work? Also, in order to not make you think I have the research to back up my thoughts, I need to say that in almost every sentence. I also don't think that should be the way hn works.
Nothing is wrong with it necessarily --- it's all a matter of implementation and audience. I think the first thing you are going to run into is a need to visually differentiate a non-vote from a zero. I'm also not sure what problem it's trying to solve.
What I would find more useful (from a 'build a better recommendations engine' perspective) is a 5+: a short list of favorites that can stand in for someone's favorites. Personally, I'd also like a better way to better differentiate the gradations between standard, good, and great. Whether I hate something or 'hate-hate' it isn't going to make much of a difference. Do you think your audience is going to be persuaded to reduce their average rating by a point, or are you still going to find the oft-quoted 4.3 average? I'm doubtful, but this doesn't make it a a bad idea to try.
As to the downvote, I stand by it. My goal is to rearrange the page so that the comments that are most useful to me are at the top. If others find your comment useful, they will see the injustice and bring it back up to the top.
As to the need for 'strong backing', I think we just have different worldviews. With due respect to my friends who are psychology professors, "a [nameless] psychology professor told me" is barely a step up from "I'm not a doctor but I play one on TV". We obviously respect different authorities in our lives.
It is not "necessary" because to me, if I believe in it and it's important to me, I will test it out. It doesn't matter if the idea came from Steve Jobs or Joe Doe. It's not "too helpful" because often the research being done on it is flawed, outdated, or just really not too trust-able. An example for example I remember reading some group has done research on max width of cell phones people like before feeling discomfort. The Motorola design team was the first to ignore that when they designed RAZR.
Now, the new steepster slider definitely calls for it since he uses 100 points rating scale. However, you have to look at the alternative, which for a product rating is normally a 5 point clicking UI w/o a slider. Thus, you have to look at both solutions and see if this new rating system that requires the slider is "better". If the number of user ratings does not decline, it would show that users are not bother by the slider, which means the steepster slider may be "better" for your product rating. But if a typical user finds the slider too troublesome and not bother with leaving a rating, would this new system be better than a simple 5 start system with a better distribution of samples? Steepster may have logs that shows that there isn't a decline in user ratings, but that is only 1 company (sample) with unique sets of users. Before another company follows, I believe it may be wise to ponder or do some testing/research on use of sliders vs. clicks for ratings.
Why even ask people how they feel?
Depending on the content, you can analysis how they use it to get a much more accurate rating. For video: Did they watch the entire thing? Did they leave after a few seconds? Did they share it somehow?
That (slightly off-topic) being said, this looks great.