YouTube Comes To A 5-Star Realization: Its Ratings Are Useless
techcrunch.com
techcrunch.com
BTW, they really should have used a histogram here.
If I were YouTube, I would start by automatically ranking the videos by the percentage of times it was completely played but discounting times it was played completely and the page was just left open with no activity (to counteract the people who opened the page and ignored it afterwards, or left their machines idle). Discard the 5% outliers on either end of the spectrum and you should have a pretty decent idea of which videos are popular and which ones are boring.
After all, don't they care mostly about whether the videos are engaging?
It seems like there should be a ton of techniques that they should have already been using to supplement the star ratings in order to fine tune any algorithms they wanted to play with.
Well, given how badly they messed up with comments...
This is key! Almost always, only good or great products will have a bunch of ratings and a distribution like:
****************
********
****
**
**
Someone is always going to dislike just about anything. But the above distribution seems to indicate something that a lot of people genuinely like.I also look at the most useful positive and negative reviews.
Then again, Amazon doesn't account for the statistical uncertainty of votes, so it's a little odd as well.
The closest I've ever seen to the "ideal" curve is the IMDB curve right next to it.
Cheers.
The econometric results reveal that the reviews for the majority of the products have an asymmetric bimodal distribution.
etc. I thought everyone knew this, and it makes sense. You don't review something if you think it's "okay". You review it if you love it or hate it. Rottentomatoes probably has a bell curve because the professional reviewers review everything, whether they want to or not.
Also, the reviews on youtube are quite useful, and a three-or-less star average indicates that something is labeled incorrectly (for copyrighted material, tv shows, movies, videos, and the like).
Maybe for YouTube they could count opinion votes when someone watches the same video twice or uses the embed/share feature. If someone watches a video but never shares it or watches it again, then I think it's fair to say they were apathetic about it.
"Five start systems are supposed to work on the basic idea that the one and five votes will average out to a value somewhere in the middle. In this way the two, three, and four star ratings are averages based on the ratio of one star votes to five star votes."
However, that doesn't always work. People seem to be lazy and they don't want to judge the comparative value of different items.
However, TC seems to believe that the 5-star system is more poorly defined. Opinions? I've been wondering about whether to include guidelines for the different rankings, (3 stars means this, 4 stars means this, etc) or just to let people use their own definitions (since, in all likelihood, the guidelines would be ignored). I've always thought it strange when reviewers provide a little box saying "5 stars - Excellent, 4 stars - Good..." and so on, since most people should have got the idea by now.
I also want to segregate links by level - Getting Started/Beginner/Intermediate/Advanced - and was trying to think of precise definitions for each. Again, I've also thought that since people likely would ignore any such definitions, it might be better for them to use their own definitions. Theory being that if 100 people thought this link should be in the 'beginner' category, then others won't be surprised to find it there.
So, give definitions for each ranking, or let users work by their own interpretations for each ranking?
2 ideas to improve things:
Idea 1: Change the weighting of a vote based on the voter's voting habits. Eg, if a person only gives 1 and 5 star votes, decrease the weighting for their votes. I doubt I'm the first person to come up with this idea, does anyone know of any sites that implement such a scheme?
Idea 2: Users with 'editor' priveleges have the ability to move things around into their 'correct' place. This could make it more useful to have predefined definitons for each ranking and category.
No, the problem is that people can't be bothered to spend time to think about just how much they like something, compared to everything else, and that it can become uncomfortable to try and do that because you realize that your relative liking may not be neither constant nor consistent, and trying to make it so could be a lot of work (I like A a lot, but B even more... hey, C is really cool! But some things about A are better than C... now what?).
Because you're implementing a system that on paper has a lot more resolution than what you're really getting. Imagine buying a 1080p HDTV and then showing solid colors on it only - not only a waste of engineering effort but also building subsequent systems around the validity of your 5-star ratings will also be fundamentally broken.
Also, based on their data there's a very concrete reason why the 1-5 star rating is worse than the thumbs up/thumbs down. With the thumbs up/down system you have a single dimension of data ("likedness"), whereas with the 1-5 star rating system they're only getting data from people who like the video (and almost none from people who disliked it, look at the distribution). This makes the data practically useless for determining the quality and user preference for a video. Consequently ranking algorithms just won't work on the star system - the difference between video #1 and #100,000 can be an average rating of 4.8 and 4.92.
"Those are subjective and different for everyone as well"
So are movie ratings - but it's still a very useful metric to a lot of people. With a large enough sample size you get the lowest common denominator preference measure - which may be what YouTube wants.
"the problem is that people can't be bothered"
I object to the labeling of basic user behaviour as laziness or some type of stupidity. Users will behave how they behave - assigning value judgments to this behaviour just makes you a prick, and disconnects you from your users (who also happen to be your customers, yay!). If your users aren't using your system in the way you intended, you need to fix it. Trying to pawn off your responsibility in the equation as "lazy users can't be bothered" simply is a cop-out.
The difference is that the HDTV costs extra, while the rating resolution does not.
but also building subsequent systems around the validity of your 5-star ratings will also be fundamentally broken.
It depends on how you interpret and use the data.
Also, based on their data there's a very concrete reason why the 1-5 star rating is worse than the thumbs up/thumbs down. With the thumbs up/down system you have a single dimension of data ("likedness"), whereas with the 1-5 star rating system they're only getting data from people who like the video (and almost none from people who disliked it, look at the distribution). This makes the data practically useless for determining the quality and user preference for a video.
Again: With only people who like voting 5 stars and a few people who dislike voting 1 star you have the exact same information as with a "favoriting only" voting system, or as a thumbs up/down one with few people voting down (which is very likely). So how is it any more useless? The only real problem I see is that people probably interpret the ratings differently - to some, 1 star is "the worst possible vote", to others it is "one notch better than no vote".
Consequently ranking algorithms just won't work on the star system - the difference between video #1 and #100,000 can be an average rating of 4.8 and 4.92.
Which is not necessarily a problem, if you weigh the average by the number of votes. Sure, you don't want to rate a video with one 5-star rating higher than one with 5000 5-star ratings and one 1-star rating. But you don't have to.
So are movie ratings - but it's still a very useful metric to a lot of people. With a large enough sample size you get the lowest common denominator preference measure - which may be what YouTube wants.
Um, yeah. That's exactly what I said.
I object to the labeling of basic user behaviour as laziness or some type of stupidity.
You're reading something into my comment that I did not say. Actually yeah, it's lazy... and people have every right to be - they go to YouTube to be entertained, not because they're being paid for it.
People do something because there are rewards. With user-generated content, the rewards are generally immaterial: attention, status, self-expression. The only reward for thoughtful ratings is self-expression, but very limited because nobody can see your vote. Even a single-word comment is more rewarding than that. So if you want people to spend the time and effort that thoughtful fine-grained ratings require, you'd have to add artificial rewards - which is almost impossible because you can't measure how thoughtful a rating is.
The alternative it to make rating easier, so that more people will do it at the current reward level. And "easier" here means "requiring less thought".
Back when Youtube rounded the average vote result down to the nearest star, I joked that Youtube videos really only had three ratings:
Five star: New video (eventually it will accrue a vote of other than 'five' and drop to four stars) Four star: Top notch (mainly fives, a few troll one-star votes) Three star: Disgusting trash (at least many one-votes as fives)
You don't see anything less than three-star since it will have been deleted by the time you get there.
So I guess a +/- system may not necessarily be more advantageous. It all depends on how carefully and correctly you interpret the data.
I don't think changing to a thumbs up or down system would affect my experience of youtube on way or another.