Python, Machine Learning, and Language Wars. A Highly Subjective Point of View
sebastianraschka.com
sebastianraschka.com
On the other hand, I've also got a higher tolerance for things not being perfect, I can figure things out for myself (and luckily have the time do so), and I'm willing to code it up if it doesn't already exist (to a point). Naturally, that is not true for most people, and thats fine.
The author isn't willing to take the risk that Julia won't "survive", which is fair. Its definitely not complete yet, but its getting there. I am confident that it will survive (and thrive) though, and continue growing the not-insubstantial community. I have a feeling the author will find their way to Julia-land eventually, in a couple of years or so.
> I have a feeling the author will find their way to Julia-land eventually, in a couple of years or so.
I have a strong feeling that this will eventually happen :). In an ideal, less busy, world, I would love to use Julia alongside to explore and battle-test it further. Or even develop useful packages, libraries, and functions for it. The truth is, I am currently lacking the time to do that :(. I mean, Python works for me, and I am currently more into the scientific problem solving so that I don't have the time :(. When I say that Python works for me I mean that I am currently happy since it can do everything for me I need, however, this doesn't mean that Julia couldn't do certain things better ;).
Anyways, I really like your comment. I am wondering if you would be okay with it if I include it in a "Other people's experiences and opinions" section at the bottom. I think this would be extremely helpful for people who are new to the "data science field" -- my article is strongly biased towards Python as you noticed :P
"There is really nothing wrong with R"
I think there is one thing wrong with R - it's name. Pretty much impossible to quickly google for help on it.
That being said, we need to realise how young Julia actually is. Last Sunday marked six years since the first commit and the language has only been publicly available for three years. Julia is as old as Python was in 1994, or as old as Ruby was in 1998. Things are rough around the edges, but getting better. Things are missing in order for you to get productive quickly, but it is improving rapidly.
I do tell people around me that Julia is most likely the language that I love the most, I tell them about the features, my involvement, etc. But, when they say "That sounds awesome! Should I use it?", I hesitate and say, well, it depends. Are you already very productive in what you are currently using? Is speed an issue for you? If not, don't feel rushed, if you want to be a "pioneer" you certainly can be and we will welcome you. But consider your own situation before jumping ship, what do you need to be productive in your day-to-day job? However, keep an eye on Julia, because I am convinced that the cost of adaptation for you will continue to go down and if there is any language that has the chance to become "just right" for Machine Learning over the next couple of years, it is most likely going to be Julia.
Point being: C++ is not hard for scientific calculations.
C/C++ is no longer best for speed, not even close.
Now, all that matters is which language has the libraries that make it easiest to get custom code onto the GPU. Python and Lua seem to be winning there, by far.
This is interesting. How is it possible that python and lua have more efficient wrappers around GPU libraries? Also there are many GPU libraries for C/C++ too. Armadillo can use NVBLAS as a backend too. I'm not sure if I get your point of C/C++ being slow.
Google / Facebook and many other huge companies are using Theano and Torch7 in production, at scale. The ML industry has been continuously moving in this direction for years now.
On these optimized ML systems, only a tiny fraction of CPU time is spent outside of the GPU. The goal in many of these companies is to migrate all tasks that can be done on GPUs to GPUs, as soon as possible. It's far faster and more cost efficient.
I would have thought that if you were going to run prod systems in the gpu you would actually write CUDA (C++) or similar to avoid the inefficiency of the abstraction layer.
(also, this comment is bordering on the uncivil).
I can only bring the horse to the water (or Google as its sometimes called these days) :)
>Google/Facebook and many other huge companies ... You are totally wrong.
That totally settles it then thank you, who am I to argue and surely there are no ML jobs that spend time outside of GPU.
(i) how nonsensical such a comparison is
(ii) there are many algorithms for which GPU offers no speedup at all in fact the the data transfer can actually hurt. There are instances where using CPU's SIMD instructions makes more sense than GPU.
Aside, I think armadillo is pretty great, especially going back and forth with other matrix libraries, also nice wrapper around OpenBLAS which made not running something on a cluster more bearable.
The thing that I like with Armadillo is that you can just pick the BLAS implementation that is the fastest on your platform and run with it, while still being high-level compared to BLAS.
[1] http://gcdart.blogspot.co.uk/2013/06/fast-matrix-multiply-an...
Just in case this is useful info, Eigen and Blaze works the same way, Blitz++ doesn't. Blaze seems to be doing the best in the benchmarks.
> OpenBLAS is pretty much on par with Intel MKL
I wont be surprised. I have seen ATLAS beat MKL on some BLAS functions on my runs. Sparse matrix multiply is particularly badly implemented in MKL (atleast was). One could beat their sparse multiply by using multiple instance of their own sparse matrix vector multiplies.
But the reason I mentioned ICC is that there is much more to SIMD than just linear algebraic operations. Non-linear operations that come up quite frequently in stats/ML are sin, cos, exp, log, tanh etc on vectors. GCC/G++ does a decent job now (the best part is that it emits info why it wasn't able to vectorize a particular loop. This lets you restructure the loops to ai the compiler), but ICC still rules.
I was under the impression that Eigen only works with MKL?
Enables the use of external BLAS level 2 and 3 routines (currently works with Intel MKL only)
Source: http://eigen.tuxfamily.org/dox/TopicUsingIntelMKL.html
But the reason I mentioned ICC is that there is much more to SIMD than just linear algebraic operations.
Indeed. I compiled some stuff with ICC, but it did not give me an tangible improvement over the latest GCCs. But for those projects, linear operations were the vast majority.
Re GCC, the same here. ICC is mighty expensive.
... that makes us co-sufferers then, or brothers in masochism :) if you will
The main reason I went for Python is purely practical: it's a language people outside my team will respect and deal with. It makes it easier for me to collaborate in many different ways: share tools with other teams, transfer ownership of my code, get help when I need it, etc. Data science at some companies has the reputation of "hack something together and throw it over the wall for someone else to deal with". In my experience R only furthers this reputation. Which is too bad, it's really great at what it does.
I like well-established languages with a large user base.
So, I was dismayed by Big Data Genomics' ADAM Project's choice of Scala, which has almost no uptake in the genomics/bioinformatics community.
They do it because they run over Spark. But Spark has an excellent Python binding.
Computation has grown more complicated. They need real computer scientists and a real language that supports real development, not some scientists which learned just enough Python to automate running some 20 year old Fortran code.
Yep. This is exactly why I gravitated toward python and scikit-learn. So much of data science is just getting the data into the right format. Grab it from this file, and this database, and this web service, then get it formatted into this table structure, and then clean it with this filter, and then plug these holes this way and those other holes that way, and now you're finally ready for a random forest baseline.
Python is a really good language for merging and parsing data from lots of different sources. For many problems, a general purpose language with very good data science library may actually be the better choice than a dedicated data science language/environment.
R is useful because there are a lot of resources as it has been along for so long and is used by a large portion of the stats community. It also has a lot of useful libraries that have not been ported over to other languages yet (ggmap!!!). But you still still run into the same problem that you cannot integrate R into a production web application.
I am pretty sure Hadoop streaming does not support R,Octave, or Matlab either
Octave/Matlab are "great" but good luck trying to integrate them
into a production web application
What problems are you facing with Octave? It has, in fact, been integrated into a couple of production web applications I know of:https://www.rollapp.com/app/octave
I have promised a while ago to improve its Python integration so that Python and Octave can be in the same process (there are lots of advantages to that kind of tight integration instead of relying on parsing output through pipes). Perhaps that could help you?
Maybe it's a good idea to implement such thing.
But I agree that integrating R with something else is an unexplored terrain. It seems that R rather tends to be everything: from data acquisition to visualisation.
Data visualizations with R seem vastly superior, unless I am missing something with Python (highly likely). And putting up a slick statistics app is easy with shiny or RStudio Presenter. But R can't really scale to a large production app, isn't that right?
So I feel I need to keep working with both Python and R.
Added: That's a nice list Lofkin. Thanks. Also, in the article he says that Python syntax feels more natural, which I also felt. But then I started to use things like the magrittr and dplyr packages in R which gives you nice things like pipes and that feeling starts to ebb.
For stats plotting and web apps in python: https://github.com/bokeh/bokeh
For calling r libraries in python: https://pypi.python.org/pypi/rpy2
For out of core datasets in python: https://github.com/blaze/dask https://github.com/blaze/blaze
Seaborn: statistical data visualization: http://stanford.edu/~mwaskom/software/seaborn/
Packages like {ggplot2,rvest,dplyr,devtools, etc.} are basically creating a sub-language for R.
I use both at the moment, but I echo the OP's ideas that R's target audience is statisticians, where Python's target audience is broader and includes statisticians and computer scientists. And Python's syntax is nicer to work with. That's why it's become the primary glue language.
That said, the overhead for learning Python as your first data science language is a bit problematic for me, as you basically have to learn Python followed by Python's data science tools (pandas, matplotlib, etc). whereas with R, you're learning the language and the data science tools at the same time, even if they're a bit idiosyncratic.
That's true - many day-to-day tasks in bioinformatics are more or less plain-text parsing [1], and Perl excels in parsing text and quickly using regular expressions. "My" generation of bioinformaticians doing data cleanup and analysis (20-30) uses Python, sometimes because plotting is nicer, the language is easier to get into, it's more commonly taught in universities, or other reasons - people older than that normally use Perl.
Both BioPython and BioPerl are extremely useful.
[1] Relevant quote from Robert Edgar: "Biology = strcomp()" from https://robertedgar.wordpress.com/2010/05/04/an-unemployed-g...
But yes, the point of that course is to implement and play around with small numerical algorithms, whereas the linked blog is about someone who mainly calls existing machine learning libraries from Python.
For someone who has never programmed before, Matlab may be more
intuitive.
Yes, this is Matlab's target audience: programmers who will not call themselves "programmers". In recent versions, they have tried even harder to hide the code away from the user, by trying to make everything work by clicking on buttons. I have heard from many Matlab users call themselves "not a programmer". They don't feel like writing software is what they're doing when they're using Matlab.I'm surprised that people are so surprised that you can pass function handles around, since there are some Matlab functions for solving ODEs or root-finding that are very commonly taught in intro courses that require function handles. I guess not everyone's first exposure to Matlab involves solving ODEs or root-finding.
Also at least in ALGOL 68, COBOL, and FORTRAN 77. Maybe earlier, too.
Try the GUI again. Here is a web version of it that is almost identical to running it on your own desktop:
https://www.rollapp.com/app/octave
This web version is based on Octave 3.8.1, though. Octave 4.0 has a much better version of the GUI:
https://en.wikipedia.org/wiki/GNU_Octave#/media/File:Octave-...
If you have Windows, try our installer:
The Matlab one-file-per-function thing, the lack of namespaces, and general lack of code structuring primitives makes it much less pleasant than Python for programs bigger than about 100 LOC though.
Dealing with higher-dimensional arrays, more sophisticated plotting, data munging, string processing, interfaces with external systems, etc. all left me banging my head in Matlab though, whereas Python makes it all a breeze.
Numpy’s broadcasting feature is also super nice, compared to wrapping everything in bsxfun calls in Matlab.
I wonder how much the @ operator in Python 3.5 will help students. Hopefully numpy can deprecate and phase out their "Matrix" object, and end the confusion about the meaning of basic operators.
> The Matlab one-file-per-function thing
By the way, Octave does not have that limitation.
I think this is actually one of MATLAB's biggest flaws. Without a true 1D array like numpy has, there is no way in MATLAB to tell the difference between a 1D sequence of values, and a 2D sequence of values with only one value along one of the dimensions.
This has led function developers to try to guess. But they guess inconsistently. Some functions treat row and column vectors differently, some treat them the same. Of those that treat them the same, some return them with the same orientation, while other force a particular orientation. Some operations ignore dimensions (length), others don't (for loops). Some maintain dimensions (size), some don't ([:]).
So everything may seem to work, until your code that has been working fine for years suddenly breaks, and you realize it is choking up because one of your experiments has only one trial, or one of your experiments has multiple trials each with one result, and some of the functions you are using start reacting differently to this. Then you have to go through each function and figure out on a case-by-case basis how it handles row and column vectors.
Or worse yet, it seems to run fine, but is silently doing the wrong thing. Which you probably would never know, because most MATLAB code isn't unit-tested.
They tolerate most languages, but find R's syntax a bit unnatural, Matlab lacking when trying to go beyond pure matrix stuff, and are waiting to see if Julia picks up (which it seems to be from what I can tell)
Also I really see Jupyter as a new standard for communication. Your narrative and supporting code all in one place, ready for sharing.
Also those classes that chose R, from my experiences, are non CS classes, the professor are from other discipline. They just want a tool that solve their need quick. An example is the Princeton's Stat class, the professor is a humanity major. The class gave us tons of data and we had to do ANOVA and such and we needed a computer to crunch so number can't do by hand. So he chose R which he uses a lot.
Which leads to another advantage of Matlab (at least as long as you don't have to pay for it) -- the documentation (including toolbox documentation) is way ahead of any competition.
Whoever came up with the R package documentation standard has done that language a great disservice :( IMO, the fact that every R package includes a huge pdf with alphabetical listing of functions and data sets is actively harmful: without it, perhaps more package authors would at least feel compelled to write a 2-page readme.txt (just like various matlab package writers do), and that would have been actually useful.