The Pattern-Seeking Fallacy
blog.asmartbear.com
blog.asmartbear.com
Most of the time, at least when it comes to sportscasting, picking the number 7 implies that's the total number of batters faced up to that point. You don't say a batter has gone 2-for-3 through 6 innings when they really went 2-for-4. When a guy is doing baseball play-by-play he doesn't say the "last 5 of 7 batters" and mean there were more than seven total and he's just choosing not to tell you about them.
Look at the kid who had a hit in every college baseball game he played this season. His hit streak was 56 games long before his team was eliminated from the post-season. If his streak ends during the first game of next season, they won't go on Sportscenter the next day and say he's gotten a hit in 56 of his last 56 games and try to hide the fact that he didn't get a hit in game #57.
Last night, during the Celtics-Lakers, game they mentioned that Ray Allen was 0-16 on three point attempts since his last make and that he was cold. No one would sanely argue he wasn't having a tough time shooting but I'm sure a math nerd could explain this away with a large enough sample size of shot attempts over a longer stretch of time then one playoff series.
But the point of a streak is that it's based on a finite set of recent attempts. No one is saying Ray Allen is statistically a bad shooter (he's actually one of the best in the history of the game) but you can't argue that in recent games he's shooting poorly from distance. It doesn't get much worse than making 0% of your shots.
That's exactly what I would argue. I think the math nerd is right. Given consistent shot mechanics, we'd expect occasional "streaks". Unless you're proposing that circumstances may change the shooter's mechanics, I don't get what you're saying.
You don't say a batter has gone 2-for-3 through 6 innings when they really went 2-for-4
You're right that they don't do that sort of thing about a single game, but they say stuff like "last 5 of 7 batters" to refer to streaks across games all the time, which makes the example I quoted a straw man.
The idea that you "expect" something doesn't diminish its importance; especially when the Celtics need Ray Allen to hit threes to improve their chances of winning. They don't care that he's hit them in the past because previous scores don't count in current games.
You think circumstances don't change the shooter's mechanics?
Consider that if a basketball player had perfectly consistent shot mechanics, and always shot from the same point on the court, he would always score. Shooting a basketball is a fairly deterministic ballistics problem--no wind since it's indoors, large target, large projectile. Thankfully, basketball gives us exactly that situation: free throws. The very best free throw shooters in the game only make 90% of free throws: http://www.basketball-reference.com/leaders/ft_pct_career.ht...
A basketball player who makes 45% of three point shots isn't some sort of machine which has a 45% probability of sinking any given three-pointer, he's a human being who varies in terms of age, physical condition, fatigue, and even psychology. When betting on black loses at the roulette wheel 16 times in a row that's pure math--when a human being misses 16 straight three point shots, it's more likely he's mentally and physically tired and his mechanics are shot.
To settle the question, you have to collect more data and carefully analyze statistics. And when you do that, the data doesn't support the existence of streaks. But you don't believe those statistics. How to convince you?
Well under your theory that there are streaks, you should expect to see record streaks out there that are longer than statistics would suggest possible. For instance if there are 500 players across time who've had a .45 field goal percentage, who've averaged 20,000 shots each, statistics says that the record streak of consecutive field goals made should be somewhere in the range 18-23 or so. If those people have even a slight tendency towards streaks on top of those percentages, you'd expect the length of the top streaks to get higher. If, during a streak, they shoot about .55, then under the same assumptions the top streak should be 25-30 or so.
So we have a test. Look at the longest streaks and see if they are longer than expected. Not even a lot longer, just slightly longer will do. Lots of sports keep track of lots of player stats, and streaks. In a fascinating essay many years ago Stephen J. Gould pointed out that, across all sports tracked and all types of streaks records are kept for, all but one are within a statistically plausible range given the number of players, games, and their playing averages.
The lone exception is Joe DiMaggio's 56 game hitting streak in 1941. That one is truly exceptional, a repeat is unlikely in the next thousand years. However when you look at the second, third, etc performances, they are in line with what statistics suggests is likely.
Now what about this college kid that you were talking about? The answer there is interesting. College is played at a more varied level than the professional leagues. Therefore there is more variation in player performance. Exceptional performances become much more likely if you add even a small number of players who are hitting somewhat more reliably. So his streak is not as impressive as it may seem.
Incidentally this applies to DiMaggio. When he played, pro baseball wasn't as consistent as it is today, so he was able to sustain stats that were above what you find in today's players. In fact http://online.wsj.com/article/SB1000142405297020455680457426... reports on a Monte Carlo simulation of baseball that used every historical player and season with their actual hitting percentages. Fully 42% of the simulated histories had a streak as long as DiMaggio's or longer.
Therefore even a slight tendency towards streaks should show up clearly in the record books. But nowhere in the record books is there any streak that shows strong evidence of streakiness.
"Instead of running multiple AdWords variants each against multiple landing page variants each feeding a different website funnel, run just one experiment at a time, one variable at a time." I think this is bad advice. If there's an interaction between your variables, this will lead you to totally miss it. Even without interactions, it can be a more efficient use of resources to estimate the effects of multiple variables at a time: http://en.wikipedia.org/wiki/Factorial_experiment
Also, in this context, it seems worth looking at http://www.stat.columbia.edu/~cook/movabletype/archives/2008...
(I also enjoyed seeing it quoted here grammatically correctly, the word data is plural, datum being the singular. Wikipedia has it grammatically incorrect but as Coase never published the phrase who knows if his grammar was as good as his maxim).
>USAGE. "Data" leads a life of its own quite independent of datum, of which it was originally the plural. It occurs in two constructions: as a plural noun (like earnings), taking a plural verb and plural modifiers (as these, many, a few) but not cardinal numbers, and serving as a referent for plural pronouns; and as an abstract mass noun (like information), taking a singular verb and singular modifiers (as this, much, little), and being referred to by a singular pronoun. Both constructions are standard. The plural construction is more common in print, perhaps because the house style of some publishers mandates it.
Yes, I am surprised. The article makes good points but its fallacy is applying mathematical abstractions to the real world. Now, there are real-world phenomena that are closer to the assumptions made by the dependency assumption in the coin tossing (Bernoulli) trials, e.g. rolling real dice. However, player performance is complex and does depend on previous history; for example a player who did badly in the previous game or missed an important shot will be under pressure this time, which will affect his performance.
The model in this case is too simplistic, a Markov model may be a better approach perhaps. However, estimating the Markov order of a given time series is a hard problem.
It doesn't matter that you can argue about different pressures, or about their confidence level, or anything else.
They ran the numbers. The tests came back "random". They generated definitely random data and people labeled it "streak".
The fact that you can come up with reasons why players ought to be more streaky than random trials is irrelevant. We ran the numbers, and they aren't.
>Instead of using a thesaurus to generate 10 ad variants, decide what pain-points or language you think will grab potential customers and test that theory specifically.
In other words, he says test one hypothesis rather than ten. But there's nothing wrong with testing ten variants so long as you have enough data to ensure that the chance that any of the variants produce a false positive because of statistical fluctuation is small.
Given a fixed, finite amount of data, you can do some pretty simple statistics to find out exactly how many hypothesis you can test at a given confidence level.
How do I know that 1000 coin flips is the right number? It seems like an arbitrary round number; how do I derive the right number of coin flips to get the 50% probability? Or maybe a better way of asking it would be: How many coin flips do I need in order to guarantee 50% heads with 4% deviation?
It was good to point out the fallacy. I just want the second step of how to avoid the fallacy.
It's about picking a theory, then testing that theory only, rather than trying to seek a theory in data that already exists.
At the end of the article it explicitly says this and gives three specific ways to solve the problem.
The measuring of those things at the end could very well fall into the realm of "streaks", too, unless we have good consideration of sample size. To use one of your examples:
Let's say that I do 10 coin flips to determine bias, and I see those 10 heads in a row. Now I want to test this situation further. Is my coin biased? I run the test 10 more times. What happens if I get a "streak" of 10 heads again? Have I verified my claim of bias? No! I just have not flipped my coin a sufficient number of times.
From this I see two things: One, an interesting anomaly arose that suggested I need to look into things a bit more. My theory is that the coin is biased towards heads.
Two, to solve the problem, I need to narrow down variables where I can (you mention) and perform the test enough times (you did not mention). We know that 10 is not enough. How do we know this? 1000 is arbitrary. What is the right number?
In other words, I agree with the fallacy you pointed out. The solutions are just insufficient. The fallacy can continue with the right probability of occurrences, even if someone tried to narrow down the variables.
I've been away from statistics for too long to go into specifics, but you can mathematically determine how large a sample size you need to be X percent confident that something weird is happening. I'd suggest reading http://en.wikipedia.org/wiki/Hypothesis_test for more info.
In particular, there's nothing wrong with the OP's example of picking 10 variants of an ad (i.e, testing 10 hypotheses) and finding out which one is better so long as you have enough data. It's really a matter of sample size.
Each observation can be loosely thought to lend power to whatever analysis you choose to run, even those including very large numbers of possible explanations. For instance:
You know, Rodriguez is 7 for 8 against left-handed pitchers in asymmetric ballparks when the tide is going out during El Niño.
is a valid statement if the model you're using includes factors for pitcher-handedness, ballpark design, and weather patterns. This is an extraordinarily broad model though almost every observation is at least a little bit novel at first, so you need a great deal of power to distinguish truly interesting things.
If you sample sufficiently (exercise left to the reader) and randomly across all those effects you'd serve a chance of learning something about how they correlate in a statistically profound fashion.
The barb of the fallacy is forgetting to watch your possibility space grow as you reach for new explanations. The moment you mention the weather you become beholden to divide your certainty by the total number of weather patterns. Unless you're very careful, that number is usually very, very large.
Randomly divide your data into two halves. Train your model on one half. Then test it on the other.
This eliminates most of the randomly discovered patterns. It also leads to a much better sense of exactly how many different ideas you tried, and therefore how high the level of statistical evidence needs to be on your successes.
This is one reason there is very little progress made on rare diseases and conditions.
Of course the same problem occurs when something is extremely common.
Anybody who has ever played sports in any sort of competitive environment has had moments when they are "on", whether it's basketball, golf, curling, whatever. Movement becomes easier. Shooting feels effortless. The game "slows down". The converse also happens, when everything feels labored and overwhelming.
Furthermore, the analysis focuses on "streak" in a very literal sense, as in runs of the same outcome that end when another outcome occurs, but the term used by sports fans often applies in a broader sense. One typically refers to streaky players as those who tend to be either very productive or very unproductive for extended periods of time, such as a game or a half, or maybe even over the course of several games.
Yes, I know how it feels. I also know the tricks that memory can play. I also know how much we tend to see patterns. And I know about statistics.
Of all of those, I trust statistics the most.
I was actually talking about colors in roulette, but the coin flip is a pretty close example. This seems extremely logical when you think about it, but a lot of people don't look at it this way. They always say something like, "Well, black has been coming up a lot lately, so it's definitely going to be red soon."
for every 20 possible data sets you cover in your trawl, you should expect to find about 1 "statistically significant" correlation that's actually a fluke (at 95% confidence, using naive statistics). If you're digging through a massive data set looking for correlations, you need to account for that, and use more advanced statistical methods.
Sportscasters often do exactly the opposite -- they quote results that aren't even particularly statistically significant (say, a 30% shooter making 3-of-4) and act as though it shows a player is "playing well" and should be trusted to continue the trend.