Mistake #0: not looking at the data.
Whether you're designing a processing pipeline or simply keeping track of what one is doing, decent visualization of inputs and outputs is critical to understanding what's going on and ensuring that nothing surprising is happening.
I'm always surprised when people come to me with data analysis problems (big or otherwise) and have never bothered to do any kind of visualization. Even if you're just visualizing sparse samples it's remarkable how easy it can be to spot issues.
Good visualization won't solve all your problems but there is a significant sub-set that go from hard to easy when you do it.