Well, O(100GB) is where I start thinking of "big data". Processing 100 GB is slow in a single "normal" machine (depending on what, sure). Between 10 and 100 I usually do it in a cluster since it is less hassle (even if I could process it locally with some tweaks or patience). Less than 10 is usually locally run unless I'm already computing something else