A data science investigation of IMDB, Rotten Tomatoes, Metacritic and Fandango
medium.freecodecamp.org
medium.freecodecamp.org
First: I claim that what we seek in a movie rating is information about whether we will like the movie, and that this can be formalized as the expected KL-divergence (information gain) between the Bayesian posterior distribution (probability of enjoying the movie conditional on its rating) and the prior distribution (probability you would enjoy a randomly selected movie). Of course, this will depend on your taste in movies, especially how much it correlates with others. But, we can _bound_ it by taking the Shannon entropy of the rating distribution: there is no way we can get more information from a rating than this! It is this bound that allows us to penalize the distributions that are heavily biased towards one side of a discrete scale, like Fandango. However, the "ideal" shape in this context is far from a Gaussian - it is uniform! The uniform distribution can also be justified as being calibrated such that the quantile function is linear - a score of 90/100 from a uniform distribution means a 90th-%ile movie. Determining a quantile is often a transform we try to perform intuitively on ratings so such a transform being trivial seems useful.
Second: The Gaussian distribution does not have bounded support! That is, a rating scheme with what you claim as the "ideal" distribution would have _some_ ratings with values that are negative or otherwise "off the scale". Not so ideal! If you wanted to model movie-goodness on an unbounded scale such that a Gaussian would have sense, then you should transform that scale into a bounded scale, eg with a logistic function, yielding an "ideal" shape of a logitnormal distribution, which incidentally can fit the strange bimodal Tomatometer distribution quite well. Even if you specifically wanted a unimodal, bell-shaped distribution, at least pick a bounded one like the beta distribution.
Third: setting aside which distribution you want to penalize distance from or why, dividing the space into three arbitrary intervals to facilitate the comparison seems ridiculous. There is already a perfectly good metric on probability distributions, the mutual information.
http://blog.moertel.com/posts/2006-01-17-mining-gold-from-th...
This was about a decade ago, so I'd expect the resulting decoder ring to be somewhat miscalibrated for today's movie ratings. But the same process would be straightforward to apply to a more up-to-date data set of ratings.
The Four Point Scale (http://tvtropes.org/pmwiki/pmwiki.php/Main/FourPointScale), while a problem from a utilitarian point of view, is still practical from a consumer psychology point of view, which is why the popular ratings systems won't change easily.
The second question I would ask is whether or not there is a relatively simple transform that could make the IMDB and maybe even the Fandango scores more uniform in their distribution, over the same set if movies?
That said, I've always preferred metacritic's scores over the others.
It's no coincidence the accuracy of ratings sites have deteriorated over the past few years, one of the most glaring violations of trust has been Rotten Tomatoes pre-certifying movies as "Fresh" and keeping it so despite aggregated reviews which would contradict a "Fresh" rating. And the even more troubling trend of "Sponsored" movies like Step receiving the certification.
It should come as no surprise the number of newly released Certified Fresh Films is increasing despite the quality of films at the box office decreasing
> Movies and TV shows are Certified Fresh with a steady Tomatometer of 75% or higher after a set amount of reviews (80 for wide-release movies, 40 for limited-release movies, 20 for TV shows), including 5 reviews from Top Critics.
Two glaring examples of "certified fresh" films loathed by audiences and top critics are Indiana Jones and the Crystal Skulls and the new Ghostbusters. Both have Audience scores in the low 50s and top critics scores in the low 60s, but "All Critics" scores are overwhelmingly positive enough to boost the score to just over 75% making them "Certified Fresh".
As far as "pre-certified" fresh, Rotten Tomatoes has taken upon themselves to sell Sponsored content and of those I've seen, Step and The Tick, as of right now, when the subject becomes Sponsored a wave of positive reviews is sure to follow.
https://www.rottentomatoes.com/browse/upcoming?minTomato=0&m...
Then what makes you think they cherry-picked their reviews? The Indiana Jones film has 260 reviews, and Ghostbusters has 325, both of which seem to be in the normal ballpark for huge wide releases. Indiana Jones is still at 77%, so it would still qualify as Certified Fresh today, while Ghostbusters is at 73%, just barely below the threshold.
I'm not seeing any reason to suspect foul play, and I don't understand why you would be upset. Audience scores and Top Critics are great pieces of information, but they're not how the Tomatometer or the Certified Fresh label work.
>Every day, a half-dozen Rotten Tomatoes staffers scour the web to find every review of every movie, collecting from major news outlets and well-known critics. They read each review and determine whether it is mostly positive or mostly negative.
>About half of the critics who appear on Rotten Tomatoes — often the more obscure set — submit their reviews, along with the ratings, to the site themselves. As reviews are indexed, Rotten Tomatoes calculates the score. http://www.timescolonist.com/entertainment/movies/rotten-tom...
The notion they cherry picked reviews comes from the more than standard deviation away from Audience and Top Critic Reviews.
As far as your last statement, it does not appear you actually understand how a review is determined to be Fresh or Not for the Tomatometer
I usually first go to Wikipedia and get a summary of the critical reviews, length of film, some other stats. Head over to IMDB for a synopsis of the film (wikipedia doesn't do summary very well, focusing instead on entire plot). Maybe skim over a couple of user reviews (both positive and negative).
If it looks interesting, I'll head over to Youtube and watch the trailer. It's always the trailer that decides for me. Having watched plenty of films over the years, I pick up a tremendous amount of info from a 2-3 minute trailer.
BTW this is my shared ranking: https://docs.google.com/spreadsheets/d/1ojCTmnu8-uIXxnas142M...
EDIT: changed 7.6 to 76 based on the comment below.
For myself, it's fairly rare that I find myself way off from the critical consensus. I'm more likely to not care for big box office action films but, then, these aren't usually at the top of critics lists either.
Recommendations is a really tough problem even within a fairly narrow problem domain like film or music. I find Amazon and Netflix to mostly be pretty bad and you can be sure they've invested heavily.
Tastes do vary; I think it's fair to say that the average theatergoer has somewhat different tastes than I do.
But. If I read reviews of the critics at the major pubs and go to the movies they recommend, I may see some films that aren't to my tastes--just the sort of thing I KNOW isn't going to be my bag going in--and I may miss a few I'd really enjoy. But it's not a bad filter. And many of the people I know have relatively similar tastes.
Maybe that just means I'm of a similar demographic and have similar tastes to many critics. There does seem to be some consensus though.
I like it a lot and have found it valuable in figuring which movies to watch in my limited free time.
Note though that I'm not talking about the engine itself recommending movies. It places you within a group and shows you the scoring data of like-minded reviewers.
The only rub in doing this would have been deciding on what an "optimal" value for the standard deviation in a 10 point rating system is, which would have been an interesting discussion. To me, this is the most important question of this whole approach. If really big standard deviations are fine, then Rotten Tomatoes uniform distribution might turn out to be the most "normal". But he seems to have totally glossed over this important fact, which is the real difference between IMDB and the other systems. With IMBD, the standard deviation of the scores is small compared to the range of possible scores. As a result, if a movie is rated 9.0, you know its gonna be pretty damn great, and a high 9 or 10 would suggest this is one of the greatest movies ever made by some distance. That's the kind of information you can't get on a rating system where 5% of movies get 5 stars.