Let’s Talk Your Data, Not Big Data
wired.com
wired.com
- (Me) So, how much data are we talking about?
- (them) 50GB
- Per hour?
- No thats our dataset so far.
- (pause)
- (pause)
(FIN)I was recently in a back-and-forth on twitter about this. Some people argued that "big data" refers to the complexity of the analysis or the value of the insight, rather than the size of the data.
Kaggle CEO Anthony Goldbloom advocated for a definition "too big to fit in an excel spreadsheet."
I advocated for for a definition "large enough that the storage and manipulation becomes part of the challenge (in addition to the analysis).
The phrase has taken on so many definitions as to become meaningless.
"Big data is when the size of the data becomes part of the problem."
Mike Loukides (O'Reilly) mentioned to me that Roger Magoulas (O'Reilly) was the first he heard using that definition.
According to this definition, physicists in the 80's were doing big data.
O'Reilly's just pimpin' the term "Big Data" to sell books. They did the same with "Web 2.0", Java and the initial internet boom.
Another part of "big data": the ability to gather, clean, and organize it in a normalized way.
-- Also: [acquisition] (not/just storage). the genius of Google and FB. they [create] massive, usable data sets.
The presumption is that big companies have a lot of data, but not necessarily. They may just have unique data needs.
Big data seems like a very similar story to me. It has nothing to do with the size of the data, not even with the complexity of the analysis, it simply means, "a movement that wants to do more with more kinds of data." Nowhere near what big data used to mean, but outside of technical circles, we're beyond the point where that even matters anymore.
P.S. If I could down vote twice, I'd give it a second one for turning a sentence into an article.
Is a TB really considered Big currently?
I maintain a database that's over 1TB and Oracle handles it very well. The trick there is understanding that, under absolutely no circumstance should you ever need to do a full table scan, because the table isn't designed for that.
So, I'd argue that a monolithic database is okay even up to 10TB even with slower disks as long as you never need to touch more than 10% of it. If you need to touch 100% of the data 100% of the time, I'd say anything over 100GB is too big for one machine.
The reality is that it just depends. There's times where you're going to want a hadoop cluster even for only 16GB of data, and there's going to be times where a database is going to be fine with 10TB of data.
(Originally from this Wavii Engineering blog post http://blog.wavii.com/2011/12/29/your-mileage-may-vary/ )
Data being generated non stop at a high enough rate that it doesn't make sense to store it. You can only analyse, extra relevant statistics or some features and move on.
Storing it is just putting in a huge buffer and as new data comes in the old data falls of the end.
In some situations where products are up 24/7 in multiple time zones, there is no time for offline batch processing. By the time the batch has finished there is newer possibly bigger batch and so on.
Note (as the example in the Wired article indicates) that the converse isn't true: just because you are using multiple machines to process data doesn't mean it's big.