For big data, what nuts does it turn? I'm still waiting to hear just what nuts people want turned, especially those for which 'big data' is essential.
It's easy enough to find cases where analysis has been stopped due to far too little data or far to little ability to handle more data. A classic example is R. Bellman's "curse of dimensionality" especially for his work in dynamic programming -- i.e., best decision making over time under uncertainty (with flavor quite different from the uses of dynamic programming in some computer science algorithms).
Broadly, for the curse of dimensionality, we can start with the set of real numbers R, a positive integer n, and the real n-dimensional space R^n, that is, just the set of all n-tuples of real numbers. Then as n starts to grow, it takes 'big data' to start to 'fill', say, the n-dimensional 'cube' with each side 100 units long, [0, 100]^n. So, if want to describe something in such a cube and want a lot of accuracy, then can start with 1 MB of data and start multiplying by factors of 1000 over and over. We can zip past a warehouse full of 4 TB disk drives in a big hurry.
Here is a general situation in 'analysis': We are looking for the value of some variable Y. Since we don't know Y, we can say we are looking for the value of a random variable. Or we can look for its distribution. For input, maybe we have many pairs (x, y) where x is in R^n and y is in R. Then, maybe we are told that in our case we have some X in R^n and want the corresponding Y or its distribution.
Well, essentially there is one, just one, one to rule all the rest, way to answer this. May I have the envelope, please? Yes, here it is (drum roll): Simple, plain old cross-tabulation. Why? Because cross tabulation is just the discrete version of the joint distribution from which we use the classic Radon-Nikodym theorem (Rudin, 'Real and Complex Analysis') to say that the best answer we can get (non-linear least squares) is just the conditional distribution or its expectation the conditional expectation, that is, E[Y|X] which is the best non-linear least squares approximation of Y as a function of X, and taking an average from a cross tabulation is the discrete approximation of this. For a good approximation for a lot of values of X, we can suck up 'big data' and ask for many factors of thousands of times more. The conditional distribution of Y given X is P(Y <= y|X) and is addressed similarly. So, net, I agree that there can be some good uses for big data.
Still, before proposing an answer and picking the tools, let's hear the real questions. Okay?
Why? For one, as general as cross tabulation is, it commonly requires so much data that even realistic versions of 'big data' are way too small. So, typically we use methods other than just cross tabulation to make better use of our limited TBs of data, and to select such methods we really need to hear the question first.
Before I select a wrench, I want to look at the nut. Is this point too much to ask in the discussion of 'big data'?
I will end with one more: Suppose we want to estimate E[Y] by taking an average of n 'samples'. Under the usual assumptions, the standard deviation of our estimate goes down like 1 over the square root of n. So, to get the standard deviation 10 times smaller, we need n to be 100 times bigger. So, roughly for each additional significant digit we want in the estimate, we need another 100 times as much data. Once we start asking for more than, say, five more significant digits, we are way up on a parabola in the amount of data we need. Net, if we want really accurate estimates, then even big data has to struggle. So, really, we accept the law of diminishing returns and just use medium or small data.