Mistakes Developers Make When Using Python for Big Data Analytics
airpair.com
airpair.com
1. Convincing themselves they have "big data" instead of just "data". If it fits in RAM on your laptop, it's definitely not "big".
2. Thinking an example of "sophisticated analytics" is "Average over time with no standard deviation".
3. Resume Driven Development over pragmatic solutions that get stuff done.
The problem I have though is that looking around for jobs recently, people want "experience in MongoDB", rather than "ability to recognize when MongoDB is clearly not the best tool for the job".
By any reasonable definition of 'big' even a gigabyte or two counts as big. And since 32bit isn't quite one for the history books just yet, many of us still have 2-4GB fresh in our heads as a threshold where you might have to start thinking about the data needing special treatment.
"Big" is always a relative term. Anything is big in the right perspective, so to say that something "isn't big" is almost always wrong. We should really try to find a term with a more absolute meaning.
Still pretty vague but it captures what I think many people are getting at.
We're sitting on several PB but it's generally in two dimensions - standard analysis techniques work, we just have problem moving the data around efficiently and processing it.
I like your inclusion of velocity, I'm including that in my definition.
Not that I've got anything against container ships, mind you, it's just that I find they're a rather unwieldy vehicle to use for getting groceries home from the market.
Nowadays a million rows is "big data" and only a "data scientist" can handle it. These guys are a joke.
We churn out many times that a day and I don't consider us to be "big data".
I'm not sure MySQL and Postgres are faster enough to justify the extra setup work when everything fits in memory. My intuition is that they may even be slower, because they still have to write everything to disk, where SQLite can be memory only.
Whether you're designing a processing pipeline or simply keeping track of what one is doing, decent visualization of inputs and outputs is critical to understanding what's going on and ensuring that nothing surprising is happening.
I'm always surprised when people come to me with data analysis problems (big or otherwise) and have never bothered to do any kind of visualization. Even if you're just visualizing sparse samples it's remarkable how easy it can be to spot issues.
Good visualization won't solve all your problems but there is a significant sub-set that go from hard to easy when you do it.
IMHO, there are too many 'data scientists' nowadays that are taking averages and calling it 'analytics'. If the call you are making exists in a library, more than likely it's not "sophisticated".
What a disparaging comment. The whole point of good software design is many of the algorithms you may need to use are packaged up for ease of use. The best ones are highly specified by the parameters you supply. "Sophisticated" should play no role in a data scientists workday - results should be verifiable and understandable, and any data analysis pipelines should be extensible and repeatable. Writing their own deep-learning implementation does not a good data scientist make.
The challenge in the field is very few folks know programming, math/stats, and the domain they're working on. It's rare to even get 2 of 3.
If all that you use is Excel, then there are a lot of huge datasets out there.
They seem super convenient but I've found that writing a lower level analysis with independent scripts linked via make or similar saves massive amounts of time in the long run.
IE the first few steps should retrieve and process the data until it is just arrays of numbers (or whatever your actual analysis needs) that can be handled with efficient numpy code.
You can use pandas for this but a real database works too. After this step pandas becomes irrelevant and just leads to lots of unnecessary allocations from recasting data on the fly.
The problem with ipython is that it leaks like a sieve and you end up with all sorts of copies of the data, worker processes and no longer used variables which will eat all of your ram even on a big cloud instance.
It is much nicer to have each analysis script start with a clean stack and release its memory when done. Plus you can use non python utilities like grep, vowaple rabbit etc as intermediate steps.
I've found practices like these significantly lower the memory requirements of analysis and allow one to tackle bigger datasets on single machines.
Each to their own though.
This also forces me to think about the engineering implications of the analysis a bit sooner. Inevitably these projects are not one-offs, and will need to be at least repeated regularly, if not outright productionalized -- and ipynb files do not lead to production-ready code.
I also agree that using a real database is often a better option than pandas. Many folks avoid it because it's less comfortable to spin up postgres + a new DB instance than it is to sit in the comfort of IPython, but it's totally worth it. Exploring and previewing the data via SQL is just so much faster and more intuitive than spinning it around with pandas.
Pandas makes a great companion to a database. Aggregate / reduce large datasets (>RAM) efficiently in a DB, then pivot in Pandas, and perform non-SQL-support analyses in Python. Same can be said for R.
iPython is just a shell, a very good one, with options to share work for presentation and reproducibility.
They are great for some things but I think the idea they should always be used is unfortunately prevalent making them anti patterns.
Yes, this is often too true - also to include Hadoop, NoSQL, ...
It's a "drop-in replacement for Python 2's csv module which supports unicode strings without a hassle."
> Doing the task in vanilla Python does have the advantage of not needing to load the whole file in memory - however, pandas does things behind the scenes to optimize I/O and performance.
There's a neat way around this, just set iterator=True or chunksize and pandas'll return you an iterable TextFileReader object: http://pandas.pydata.org/pandas-docs/stable/io.html#io-chunk...
Another tip is that even before you get into cythonizing stuff, see how much of the computation you can push down into high performance libraries like numpy - use a map() with a vectorized numpy function, and store your stuff in a numpy data structure instead of a manual for loop on a regular array-of-arrays, etc.
Sidenote, is there a decorator version of cythonmagic? Basically I want to annotate my functions with it since sometimes the non-typed basic cython is still much faster than the pure python version, and I don't have to manage the compilation step.
Honestly though, the biggest problem with people using python for bid data though is just the opposite of number 2, sometimes you need to use that framework, but when you rely too heavily on various frameworks, that's how you end up in dependency hell, largely due to the aforementioned format issues. eg. Framework 1 update fixes bug X but creats bug y in framework 2's parsing.
As far as your second point though: although it's not ideal, don't virtualenvs do the trick? If you really needed to, you could even set up a workflow of different sandboxes to pass data through. If the alternative to relying on frameworks is rolling your own, frankly I would probably choose the former, but that's just me.
The other thing to remember about roll your own is you can get efficiency that frameworks can't, and that time difference adds up fast. For example, we designed a worker distribution system, rolled our own worker management code on top of a framework, and reduced compute times from ~1-2days to ~4 hours. That's a huge increase in productivity that no tool or framework could give us.
There is power in roll your own, I would just say whip your programmers into submission regarding good commenting/documentation though.
Not necessarily a bad post, but the "Big Data" title misleading.