How Python became the language of choice for data science
blog.mikiobraun.de
blog.mikiobraun.de
My impression on two small bits of this article:
Python is also somewhat restrictive with what you can say on a single line. In Matlab you would often load some data, start editing the functions and build you data analysis step by step, while in Python you tend to have files which you start from the command line (or at least that’s how I tend to do it).
I actually find that this kind of analysis is more conveniently done in IPython (or IJulia) Notebook than in MATLAB.
I still have dreams of a plotting library where the type of visualization is decoupled from the data (such that you can say “plot this matrix as a scatter plot, or as an image”).
This is the design goal of both R's ggplot2 and Julia's Gadfly.jl.
I think a good start for julia would be to include more workflow-oriented tutorials in the docs that also provide a path through the current mess of packages (220 at current count, with sometimes quite redundant functionality, e.g., Gadfly, Winston, pyplot, Gaston). Heck, maybe I should just help with it once I'm done with my current project.
My recommendation to other numeric Python devs is that if you don't have any investment in a specific numeric Python library think about moving your project to Julia, I suspect the ecosystem will bloom pretty rapidly and provide a much more future-proof set of tools for technical computing.
I remember the days when I tried to do command line automation in Matlab, or tried to read some XML from some internet resource.
This stuff is super simple in Python, but very hard in Matlab, because Matlab focuses so much on science and so little on general usefulness.
Python is not as integrated with regards to the scientific core functionality, but it doesn't limit you to it, either.
Matlab has come a long way though. Coming from a functional programming background, I was delighted to find that you can pass functions as arguments to some matrix operators and get some very terse map+reduce operations.
fname = "/tmp/myfile with spaces"
f = open(fname, "w")
run(`echo hi` |> f)
close(f)
pipe, process = readsfrom(`cat $fname`)
cmdoutput = readall(pipe)
println(cmdoutput)
I'd write all of my command line scripts in Julia, except the interpreter currently takes far too long to start. But people are working on that.But pure matrix math is a small part of my analyses and workflow. Most of the code is in data preprocessing from SQL servers, CSV files, JSON files, and various cleanups. Once you have your data in an array, MATLAB can be great, but writing general-purpose code (or even a while loop) was horrible.
In general, unlike many other technical computing languages, Julia does not expect programs to be written in a vectorized style for performance. http://docs.julialang.org/en/latest/manual/arrays/
In Fortan, MATLAB and Python, the opposite is true: vector notation is fastest.
Ggplot2 and friends are great, but theres plenty of room for even better tools in the data vis space.
My impression (and I am not a statistician/data scientist in my day job, so I would very much like to hear opposing perspectives) is that the R ecosystem is far more mature and widely adopted than scikit-learn for things like regression, classification, clustering etc.
The author also cites expensive MATLAB licenses as a driving force behind python adoption, but here too I'm skeptical. As a grad student, I get MATLAB for free. But I switched to R/pandas for data analysis because R has a native data structure for working with multidimensional datasets (i.e. data.frame).
To illustrate the utility of this, let's say you asked developers all over the US for their zipcode and salary and recorded the results in salary.poll.data. Here's an interesting question: what is the mean salary in each zipcode? In R, all sorts of libraries make this computation concise and highly readable. Using the excellent data.table package, you would do `salary.poll.data[, list(mean.salary.by.zip.code=mean(salary)), by=zipcode]`.
No such libraries exist in popular usage for MATLAB. You'd have to roll your own, or more likely, write a lot of crufty loops and conditional statements. (Or use higher order functions with map/reduce/filter, which, by the way, you would have to implement yourself).
For me at least, just having the right data structures for working with data makes R/pandas a clear winner for doing statistical analysis of data.
The thing that draws me towards Python is that it's a well designed, general purpose programming language. The syntax is sensible, the language allows real OO and FP abstractions, you have easy access to basic data structures like lists and hashmaps, and there's a huge ecosystem of third-party libraries to build on. Things that are stupidly difficult in MATLAB and R, like file/network IO, string processing, or building a GUI or web interface, are straightforward in Python. If you've ever tried to run MATLAB on a cluster, it's an absolute nightmare. You would never, ever think about building a production system in MATLAB or R. But these things are easy in Python. There's something that's just really nice about having access to first-class analytics tools in the same language that you're building your systems in.
Other than the last one, all of these are available in Matlab and R as well; and all three have a huge set of libraries -- the difference is about the focus of those libraries imo. Python does have more general-purpose libraries, just as R has more statistical packages.
I have certainly run a lot Matlab on a cluster, in fact, I'd say safely about 70% of code I see running on clusters around here is Matlab. I've also seen a few production systems in (mostly) R -- I actually suspect that at least on Windows, deploying an R system may be easier, since you just install R and then use internal functionality to get packages you need; python (with all necessary extensions) seems relatively tricky to get running.
All that being said, I do agree that if you are building a production system where most of the code is related to interfacing with other systems, or GUI's etc., and the data analysis is a small and non-interactive part (ie. no data exploration), Python is a very reasonable choice if you can keep the whole project in it.
"highly readable"
Are we talking about the same language? The only reason R ever saw use is it's adoption by the statistical community, but all other things considered, R is a shitty language.
-Awfully difficult syntax
-Extremely high learning curve
-Nightmare to debug (error and warning messages are the most cryptic I've ever worked with)
-Some of the worst documentation I've ever encountered
-Many transformations are not visible to the user (how special characters are handled for example)...
-The largest set of statistical libraries...but have you ever tried using any? Good luck! Next to nonexistent documentation for most, and even with a background in mathematical statistics, I still have a hard time following examples
-More often than not, lack of backwards compatibility
I'm a statistician by training and data scientists by day job, and have used R for years. I really admire what some have tried to do with the language (Hadley). But, it's a result of the statistical community being light years behind on computational training and is by no means a good language; it just happened to come along at the right time and had some really genius initial developers.
Frankly, the "language of choice for data science" simply doesn't exist yet. You just use what works best for the problem at the time, and more often than not that involves switching between many languages.
means_by_zip = grpstats(ds, {'zipcode'})
[1] http://www.mathworks.com/help/stats/dataset.html
[2] http://www.mathworks.com/help/stats/grpstats.htmlhttp://blogs.mathworks.com/loren/2013/09/10/introduction-to-...
So you don't need a stats license for that anymore. The question is how many functions will accept this datatype, which is one of the key problems with some of the more advanced datatypes (timeseries has similar issues).
I didn't have much trouble with datasets because it's fairly painless to convert to a matrix with double(ds) if you setup your text columns as ordinals and nominals rather than cell arrays.
R:
- has more ml libraries.
- is more mature than panadas (look at how pandas handles categorial data).
Python:
- has more general purpose libraries (e.g. language detection, html boilerplate extraction, query apis)
- can handle out of core learning (datasets, which don't fit into memory)
- can run in production.
- is language of choice for deep learning
Could you expand on this point?
Python scores over R in one aspect - text processing.
Python is definitely major player in data processing (heck i'd even just be interested in how he's defining data science), but it's definitely not the only game in town by a long shot.
Summary: "Constantly switching languages is a chore, and while Python isn't the best at everything, it's the best at some things and good enough at almost everything else."
As a long-time Matlab user, I'm currently working on migrating my skills to Python simply because the small amount of functionality I may lose (decent IDE with integrated debugger and workspace viewer) I gain back with loads of other tools (like extensive out-of-box support for databases, way better memory management, and exciting developments for targeting GPUs and compute clusters).
Personally I am not ready to hand the title "langauge of choice" to Python, although it is trending that way. We should give R credit where credit is due. There is still a huge population of R users and code. Python has advantages in that the syntax is easier to understand, and so is the structure (object oriented versus scripting). What Python lacks is a simple setup. I think that once Python becomes more accessible to everyone (as far as downloading Python, setting up directories, packages, libraries, etc.), it will have huge leaps in usage.
I don't actually have it myself, but it does seem to be trying to solve the problem you point out.
http://code.google.com/p/spyderlib/
we ship it with anaconda
Python, Matlab, Fortran, Cobol, will be around for a VERY long time because so many of the smartest people THINK in these languages. The number and quality of people who think in a language is more important than the number who develop in it.
I don't yet think in Python. It is not where I learned programming. I am more a Lisp thinker, but for practical application python is a better choice.
I don't trust people who think in JavaScript. Or rather I don't like to bet on them.
I also believe it's undeniable that parts of these languages would be a godsend in some more "real world" languages.
There are two sides to the argument, however I'd like to caution against dismissing languages as "not for real world use" too quickly since it's a trend I've seen.
One thing people might want to consider before investing time in Python is that I found it to be quite memory inefficient: data structures take up a lot of space, and the garbage collection didn't seem to be as effective as other languages (I spend some time studying/improving GC in JVMs). So if you're dealing with large amounts of data and/or complex data structures, I wonder if Matlab might be more appropriate (AFAIK R is also not very good at memory management yet).
In most cases running out of memory in matlab meant either making the problem smaller or running it on a beefier machine. I think this is the reason why you see a lot of labs at universities with machines that have 96GB of memory, even though their datasets seem to be much smaller.
FWIW, as far as processing lots of data is concerned, python is not without issues. If you do it naively you will run out of memory really quickly. But by picking your tools correctly you can go a long way. Use Pandas and/or Sparse arrays whenever possible. Learn how numpy broadcasting operations contribute to memory explosions. Take a gander at the source of that sklearn method you're using, since it's often quite obvious that the particular implementation will choke.
I've found that these days I try my best to avoid loading datasets into memory. This is second nature for people who work with 'big data', but it's an m.o. that takes some getting used to. That is, blocking and/or streaming your data, and appropriately subdividing your problem for distributed computation. It's worth mentioning that this problem with python is under active research and development. The guys at continuum developed IOpro to deal with the issue of memory efficiency when loading data, and to make streaming data from flat files/S3/mongodb/whatever easier and more stable. Also, their (very young) project called Blaze is meant to be a drop-in replacement for numpy, but is designed for efficiency and specifically for dealing with out-of-core computation. We'll see...
I'm not experienced enough in using Matlab, but I did see a seminar given by some of the engineers and they seemed to have given a lot of thought to optimising the software for large datasets.
> I've found that these days I try my best to avoid loading datasets into memory. This is second nature for people who work with 'big data', but it's an m.o. that takes some getting used to.
Yes, couldn't agree more. I'm still getting used to working this way.
Just for whatever it's worth, this is the whole point of numpy. Lists, dicts, etc are not memory-efficient, but numpy arrays are.
The main reason I switched from Matlab to Python is due to Matlab's excessive memory usage (Essentially every operation makes a copy). Using numpy, you have a lot more control over memory usage than you do with Matlab.
http://kmike.ru/python-data-structures/
For preprocessing you can usually work on a line-by-line basis (e.g. don't keep the entire dataset in memory).
This is a bigger problem for web apps, imo. But there you usually just throw money (e.g. more servers) at the problem.
For example, the everything is a dictionary approach that widely used NetworkX library utilizes was a horrible memory gobbler on large graphs. It sounds very nice in theory, works well on smaller graphs but is near useless for larger ones.
The best graph tool that I found for larger graphs was graph-tool which is heavily NumPy based.
When I write in C, at least I know where my memory goes.
R vs some Python alternatives in Google Trends: http://bit.ly/1fpOR57
Pandas is looking good, but then we have Julia, Matlab and probably some people still using SAS.
edit:
Comparing statistical packages by: Kaggle.com usage, Job posts, Activity on blogs, mailing lists, Stack overflow, Google Scholar and a few more..
http://r4stats.com/articles/popularity/
We have a winner?
Julia... well... has its merits, but the ecosystem isn't comparable. Julia does not have Django, Guido or PyCon.
Regarding ecosystem development, one advantage that Julia has is the low learning curve between "user" and "contributor". Idiomatic code usually runs reasonably quickly, and can be made "fast" by tweaking some things within the same language (devectorizing, or turning off bounds-checks, for example). These kinds of optimizations simply cannot be done in NumPy without dropping in to C+CPython or Cython (each of which has impedance mismatches and language+toolchain hurdles).
Sigh.
HBase? Java
Hive? Java
RapidMiner? Java
Cassandra? Java
Neo4J? Java
Python 4 the win.
Unfortunately, this means that most data science findings, when transmuted into permanent production programs, don't stay in Python. Often Java or C++ are used instead.
Personally, that's why I think Clojure's got a real shot at being "the language of choice for data science" in 2023. It has the power of Lisp, it makes a lot of data-frame manipulations really easy, and because it sits on top of the JVM, it can be "productionized" pretty easily (you might have to write a couple functions in Java, for performance).