The statistically interesting aspects come from a large number of variables, not observations.
Edit ... And so I think the comment's focus on the curse* of dimensionality and "small data" is misplaced.
* edit #2
The statistically interesting aspects come from a large number of variables, not observations.
Edit ... And so I think the comment's focus on the curse* of dimensionality and "small data" is misplaced.
* edit #2
Your point about what is 'big' about 'big' data is not really many more observations but many more variables. Okay. Of course there are problems trolling for causality or even just reliable relationships; we may find something in our net that is not real. But, if from big data we select and work with just a few variables, then with just those variables we are back to small or medium data.
So, the shortest 'point' is that it is so far not so clear just what is new and good that really needs big data. So I asked, for the big data wrench, what nut does it turn that needs turning? I'm not saying that there is no such nut; instead, so far I'm not hearing what the nut is and not seeing it on my own. So, I'm asking the big data people, what are the problems that want big data to solve. This goes to the OP claim of the importance of analysis of big data. I agree that the analysis is crucial, but without more idea of just what the problems, nuts, are we still need more to get excited about the opportunity for valuable analysis.
Obviously, none of this stuff is useful if it finds meaningless relationships; that's always true when people are looking at empirical data and is not unique to big-data. There is a lot of research on dealing with that exact issue in this setting. An old paper that looks at this stuff from an economics/finance perspective is here:
http://weber.ucsd.edu/~hwhite/pub_files/hwcv-077.pdf
and there's been a lot of research since then.
I suspect that the popularity and faddish nature of "big data" right now comes from hope that those automation procedures will make it conceptually easier to do data analysis, but I don't really know if it will (and I have some doubts).
[1] I am not a big-data person, so anyone more knowledgeable should jump in and correct me.
The problem being that as we add more data, the chances of spurious relationships increase dramatically, and the human brain is incredibly good at finding a causal explanation for those relationships, even if none exists. This can quickly turn Big Data into a noise-generating rabbit-hole, leading us down blind alleys, and wasting our time.
I love that it keeps getting easier to test our hypotheses, but a search that begins without a logical and reasoned hypothesis is a dangerous beast.
If have enough data and keep testing hypotheses long enough and keep fitting long enough, then have a good chance of finding a hypothesis can't reject and a fit that looks good, even though are looking at junk.
So, divide the data in half, fit to the first half and test the fit on the second half. And if the fit fails on the second half, then what? Return to the first half, fit again, and then test on the second half again? Now want some more data to test the most recent fit.
More can be done along these lines.
That's one thing that's interesting about this stuff from a statistics perspective, how you can draw conclusions that are reliable even after some sort of search process. See, for example, the research by Joe Romano and Michael Wolf (and their coauthors) on stuff like "family-wise error rate".
Some people (like you) understand that reliability and validity aren't just "p < 0.05", but that's far from universal understanding. I've seen intelligent people accept and reject hypotheses with woefully inadequate evidence, and I've also seen wild hypotheses built on the backs of strong but meaningless correlations.
Dangerous beasts can be useful, but they must be treated with due care.
Just trying to learn how to better read data
A customer's marketing group was tying visitor data to geodemographic data. They put together a database with tons of variables, went searching, and found a multiple regression with a Pearson coefficient of 0.8+, a low p, decided to rewrite personas, and started devising new tactics based on the discovery.
Fortunately, they briefed the CEO and the CEO said that the dimensions in question (I honestly don't remember what they were) didn't make intuitive sense, and demanded more details before supporting such a major shift in tactics. More research was done, and this time somebody remembered that this was a product where the customers aren't the users, so they need to be treated separately. And it turned out the original analysis (done without fancy analytics) was very close to correct.
If the CEO hadn't been engaged during that meeting, they would've thrown away good tactics on a simple mistake. The regression was "reliable" by most statistical measures, but it was noise.
A similar example holds for validity, where I saw a team make wonderfully accurate promotion response models, but they only measured to the first "conversion" instead of measuring LTV. And after several months of the new campaign, it turned out that the new customers had much higher churn, so they weren't nearly as valuable as the original customers.
> Care to elaborate how how to be more sure of reliability and validity?
I'm not a statistician or an actuary. I'm a guy who took four stat classes during undergrad. I know just enough to know that I don't know that much.
Disclaimer aside: my biggest rules of thumb are to make sure that you're measuring the thing you want to measure (not a substitute), to make sure the statistical methods you're using are appropriate for the data you're collecting, and to make sure you understand the segmentation of your market.
But, yeah, you hand some people a spreadsheet with numbers in it and their critical thinking ability just evaporates.
As an aside, that's not what I meant by "reliable" earlier (and, to be really specific, I agree that low p-values do not ensure reliability even w/out the other problems introduced by that particular model search).