For example, if I can have a single machine with 32 cores and 1TB memory? what is massive in this context?
For example, if I can have a single machine with 32 cores and 1TB memory? what is massive in this context?
I'd define "massive" data as anything where n^2 is too big, where "too big" is bigger than either my ram or my patience.
New issues appear when you have to analyze 2Tb with a 32gb RAM machine, but when the order of difference is the same, the issues and thus the answers are the same as before?
Also, the rest of the use cases (which fits into a single machine memory now), can be handled much more efficiently with memory base algorithm, instead of I/O based algorithms.
The goal of Hadoop, as well as most of the theory on disk-based indices (E.g. BTREE), was to overcome the I/O bottlenecks. But as memory is getting bigger and cheaper there is a trend to drop Hadoop in favor of reading data directly from the cloud and into memory.