Movie Review Aggregator Ratings Have No Relationship with Box Office Success
minimaxir.com
minimaxir.com
Indies form the second clump high in the rankings.
The negative correlation is caused, as the density contours show, by a knot of highly reviewed indies showing up in the data, confusing the positive correlation of the studio movies.
Indie cluster alone:
log-corr: -0.1243865
corr: -0.1673577
Blockbuster cluster alone:
log-corr: 0.230726
corr: 0.2570628
Maybe I should have separated the clusters, and I'll definitely do that for further analysis (even though at best, the correlation is weak in either direction, and the log doesn't change much). I did, however, want to address the general reliance on RT score for any movie, so I stand by the post.
The reliance of RT score for what? What are people routinely extrapolating from the RT score that you think is not valid?
Extrapolating budget from the RT score is definitely invalid: you've shown that. But who does that?
People might go the other way: extrapolate quality from a big budget (it must be good, it's a big blockbuster), but as you say, that correlation is valid (though noisy).
So not quite sure what you're standing by, sorry. I don't want to sound rude and argumentative. I'm not trying to be a jerk, I'm very happy to concede your analysis is applicable where it is, I just genuinely don't understand what you're referring to here.
This avenue doesn't immediately distinguish whether a movie is a blockbuster or indie or how much money it made. But then again, some people do read movie descriptions thoroughly...
So you think box office is a better correlate of 'good movie' than RT score? I didn't get that from the post, but if that's what you're saying, then yes, I concede you have shown that correlation is not valid.
I'll update the post with a footnote tomorrow clarifying things.
The reason for this is simple: movie critics are self-selecting movie aficionados. Their tastes just don't reflect the tastes of the public at large. When you aggregate critical reviews, you're only aggregating the opinions of a tiny sample of the movie-going public, and that sample largely shares the same tastes.
https://www.youtube.com/watch?v=cDmluUKLbYc http://www.theguardian.com/books/2011/sep/08/good-bad-mulitp...
https://cannes-rurban.rhcloud.com/
What is evident is that very often critics gather to champion objectively awful movies. Even for professional movie critics standards awful movies. Looks for example at the surprising success of this years "Carol" and "The Assassin" in Cannes and their disappointing acceptance in the real world. Well, "Carol" is one of those special needs cases, which still has some champions left over. And compare it to the exceptional Cannes movies this year: Dheepan, Umimachi Diary and Embrace of the Serpent. The critics didn't appreciate them and didn't see it coming. However the professional festival programmers and jury in the followup festivals saw it.
But the same happens almost every year. "ADIEU AU LANGAGE" (Godard) highest rated movie 2014, "HOLY MOTORS" (Leos Carax) 2012, "Le Havre" (Aki Kaurismaki) 2011, "FILM SOCIALISME" (Jean-Luc Godard) 2010. The jury votes are usually much better than the critics votes.
Comparing quality to quantity (advertising budget, imdb ratings) makes no sense at all.
It would be interested to see how the picture looks when controlled for movie budget, time of year released etc.
Rotten Tomatoes was the only one that had any hint of meaningfulness. The others (IMDB, MetaCritic had none). Twitter volume, sentiment and number of theaters it opened in were pretty good indicators though.
A movie will have a reputation: good, bad, or indifferent. But a movie also has to be known. Ad spend gets the movie in front of eye-balls at least insofar as it gets to be known to exist. But also word-of-mouth, which is like free ad spend. This is why there are slow burners and mega-flops.
Incidentally this is why I have refused to see Titanic, Avatar, and the latest Star Wars in the theatre. If you scream in my face I'm not going to watch your movie in the theatre. Hype turns me off seeing something, but that's just me not wanting to roll with the crowd perhaps, and not a function of hype itself. Seems like most people don't react this way though.
You've caught my interest here - Is this out of dislike for advertising as a whole? Do you apply this principle to non-movie situations? (TV, food products etc)
Next article: Evaluation of Paintings by Well-Regarded Contemporary Critics Has No Relationship to Future Dorm Room Poster Sales Figures.
Edit: Also, a measure of success that doesn't find a way to divide by the budget somewhere is not a good measure. A lot of flops have made 100 million dollars, they just cost 120 million to make.
I would like to see this analyses run on box office as a percentage of budget, not just top line revenue.
Let's say you take a friend from another country to a baseball game. He's never seen baseball and knows nothing of the game. He can certainly enjoy himself -- there will be lots of colors, activity, and it's quite a spectacle. He may even want to come back.
As he learns more and more about baseball, however, he will have a completely different experience. It's the same game, but now you see a lot more of the complexity and nuance.
As a film fan, I use the meta-ratings almost exclusively. I'll go and watch popular films. Sometimes I'll even enjoy them. But there's a lot going on in cinema. I like understanding the nuance and detail, and I think well-mace movies make for better experiences.
1. http://www.thegreatcourses.com/courses/how-to-listen-to-and-...
Sure, you will experience it differently if you know what good tactics look like but most people enjoy it because they are emotionally involved with the teams and players. This seems to be true with most sports, even for the more "knowledgeable" fans.
Outside the US it's a harder problem http://arxiv.org/abs/1405.5924
Although I said I wouldn't disclose the data, here is a spreadsheet of the Top Movies by Box Office Revenue with a RT score <= 20%: https://www.icloud.com/numbers/0005voEEmzex1t4xnMHD3u3HQ#box...
Films like Alvin and the Chipmunks and Grown Ups are objectively bad, but they still make money because people don't care. Yes, these movies likely had a lower budget, and that can be weighed against box office revenue, but I explained my concerns with that in the article.
I mean no offense to anyone. But there's a lot of folks out there living a very different lifestyle than the average HN'er.
I mean, something has got to account for the fact that Adam Sandler still has an audience.
That's why data analysis is always important, even if it's dumb.
Surely, correlation provides information on association rather than cause and effect (causation should rather be modeled with Granger and other regression models). Sample sizes and variances will certainly contribute to different p-value outcomes. This is because p-values reward low variance more than the magnitude of impact (Type I/II error etc.). If you have p-values, better report them and add a footnote on how to interpret them.
Technically, in this case, a significance test would answer the question "is the Pearson correlation statistically significant from 0?" In this case, we would expect it to fail since it clearly isn't, and is therefore the test is less helpful/important. (even if it passed, the conclusion would be "correlations are low in magnitude and therefore do not matter" as noted in the post anyways)
Finding the exact P-value of a Pearson correlation requires setting up bootstrapping, which is not something I have handy at the time but will work on in future posts.
Again, I'm not looking at R^2 and the P-value of a linear regression, which is different.
In small sample sizes, correlation can easily be significant, often at the cost of low confidence. To the opposite, in large sample sizes, the magnitude of the effect may be lower but at higher confidence. In both cases, results have to be interpreted with caution. The recent p-value debate points towards a lot of issues here. For instance, there have been medical studies overestimating correlations in small sample sizes while other authors seemed to underestimate their long-term large-sample results with correlations in the ballpark of 0.15 (p<0.05).
The long version: So, in this case, we have two variables, Metacritic score (or Rotten Tomatoes percentage) and box office gross. We have measured the correlation between them for some number N of movies. N is smaller than the total population of movies that could be evaluated; if nothing else, the analysis doesn't consider movies that haven't been released yet. So the movies evaluated are considered, for the purposes of a p-value test, to be a sample of size N from a hypothetical infinite population of movies with box office gross and Metacritic scores. The p-value test also assumes that there's something called a null hypothesis, which is that there is no relationship between Metacritic scores and box office gross. The null hypothesis is the hypothesis that all other hypothesises (is that the right plural? I don't know) are evaluated against.
What the p-value measures (in this case, it can be applied to other statistics as well) is the probability of seeing a correlation of that size or greater given a random sample of N observations if the null hypothesis is true. What's notable about this is that the p-value is not testing anything about any hypothesis other than the null hypothesis -- it can be used as evidence about the likelihood of the null hypothesis being true, but past that it has nothing to say about any OTHER hypothesis. Which is why it's wrong to say that a p-value test is evidence of a causal relationship -- it's not even trying to test that.
I just realized I misread the GP's comment: thought it said "P-value" in general instead of "P-value of Pearson correlation," as a result I thought he was referring to a regression P-value.