NOTE: there are people in the world who would laugh at my definition and say that big data starts at 1Pb.
NOTE: there are people in the world who would laugh at my definition and say that big data starts at 1Pb.
If it fits in RAM on your laptop, it isn't big data.
If you can't process/handle it in a reasonable time on a single machine and your methods need to explicitly worry about how to scale to handle the data volumes it probably is "Big Data".
Problems that are embarrassingly parallel need far more data before I'd consider them big (I'd be in the >10PB camp), whereas for relational data I'd say >1TB.
These days Amazon have Multi-AZ RDS, which should handle the 2nd item.
You can't practically fit 50TB on one machine and have reasonable performance, that means multiple machines with the data spread across them.
There's then two potential issues: 1) You're doing 1-to-1 joins across tables in a query, network latency may be an issue at high query rates 2) You're going 1-to-many or many-to-many joins across tables in a query, the resulting combinatorial explosion of data is too much to handle
You want to have your inner loops/joins as deep down in the stack as possible. If you can structure things so all the heavy lifting stays inside one rack/machine/NUMA node/ processor/core you'll be able to scale a good bit further further.
Designing things not to require joins, denormalising and putting it in a column store like Cassandra is also a good approach.
Another is not having pockets deep enough to solve it with intellectual property, either in the form of a parallel proprietary rdbs (spensive) or the need to implement clever stuff.
Big data as a technology is about dumb as brick, cheap as chips, brute force.
I'm you were trying to do anything that's O(n.log n) or O(n^2) (think graph processing) then you'll run into trouble at much smaller scales.
It's not just eh size though, IMO it implies a certain dimensionality and/or lack of structure. At work were sittin on several petabytes and I don't view it as Big Data because it's actually pretty simple. We share many of the same problems as Big Data but not all
1PB is arguably enough data to store genetic variation across all human beings.
I was working on the principle that the effective population size of humans is 10,000.
(And your genome is oversized, no? 3 billion base-pairs is less than 1 Gigabyte)
I commend them for having a larger penis ^H^H^H^H^H^H data stack than you.
I thought big data was less about the actual size of the data store and more about where it comes from (typically passive collection from user activity) and how it's accessed (through some kind of large map-reduce style framework) and used (to inform product decisions or learn more about human behavior)?
An excellent example was on HN the other day, using the NYC taxi data to determine which drivers are observant Muslims. It's not something anyone set out to record, but the data set has gotten so large that if you turn it sideways and shake, random facts like that fall out.
Big data is neither big nor particularly complex.
If it cannot, then you pay the price of all the complexity and overheads of big data processing techniques so that you can get your processing done.
It's correlated with data size, bot not so strictly - you can get, for example, NLP processing problems where you need a painful pipeline split over a huge cluster for a single gb of input data, and you can have problems where the best way to process a petabyte dataset is just to stick a single powerful machine to get the performance benefits of locality and low latency, and avoid managing splits/failed nodes/whatever.
So, in the first problem you would need to use Big Data techniques and the second problem you don't, it's not related to big data and the recommendations on how best to do that won't help people who need to do big data processing.