@ Imagine you have some numbers spread over some computers -- too many to fit in one computer find the median.
▪ Uhh, sort them?
@ Can you find the median on a single computer without sorting them.
▪ :-(
@ We'll call you back tomorrow.
I was promptly rejected, but it set the tone for my later studies.
The criterion for Big Data seems to be that it fits on thousands of computers, perhaps several TB or a PB. Then I had to think of some examples:
* A million YouTube Videos
* All the tweets in the US in the past 15 minutes
* All US tax records
I still think the map-reduce philosophy is really cool. And I know at that scale there are special counting algorithms (like Bloom Filters) that may lead to some improvements at the GB or MB scales.