Programming Languages for Machine Learning
hunch.net
hunch.net
http://www.cse.unsw.edu.au/~benl/papers/stencil/stencil-hask...
The most glaring problem is that GHC currently does not support SIMD vector instructions. A problem which is now, fortunately, being addressed:
http://hackage.haskell.org/trac/ghc/wiki/SIMD
But currently, Haskell is not really an option if you want to write serious ML software that rely on vector or matrix operations.
[1] I did, for maximum entropy modeling. Believe me, it's not near competitive to C or C++ yet. But one day, it will probably be.
In this case, you're also guilty of putting optimization before the algorithm. Squeezing out the last clock cycle doesn't matter when you have raw data in one hand and some high-level objectives in the other. Effective machine learning is more about finding ways to obtain and reduce data into a useful form, and finding ways to map between your design goals and workable algorithms. Speed is the last step, and often the only one which can be improved by throwing money at it.
Besides which, if you want to have a pissing contest about speed, C is still going to lose to FPGAs. If it's that important, hire someone who knows Verilog.
What niches? ML is exactly one of the niches where highly optimized code helps tremendously. We use ML, amongst other things, for parse disambiguation and fluency ranking. Making a modification in, say the grammar, often results in training and retraining various components that have a mutual influence (e.g. parse disambiguation, auxiliary distributions for subcategorization frames, the part of speech tagger). Evaluation is also performed using ten-fold cross validation.
Since a modification can also result in regressions, the effect of modifications are usually checked individually. Being able to validate changes in one hour rather than one day, helps the development of such a system tremendously.
We developed some software first in other languages (Java, Scala, Python), usually to make a prototype for getting grip on a problem. But reimplementing it in C or C++ paid of tremendously on each occasion. Even though there is enough CPU time available (we have a 3280 core cluster).
Another thing to note is that most machine learning techniques are fairly generic. So, usually there is already a perfectly fine C/C++ implementation available. Sometimes with bindings and all.
In this case, you're also guilty of putting optimization before the algorithm.
How do you know? Most data is actually preprocessed using Prolog or Perl. But the actual machine software is written in C or C++.
Speed is the last step, and often the only one which can be improved by throwing money at it.
Again, there are niches where it is worth it. And machine learning is often that niche, because people use huge data sets, complex problems, or both.
Besides which, if you want to have a pissing contest about speed, C is still going to lose to FPGAs.
Yes, but FPGAs are not readily available to our users :). But we are very much interested in GPU computing.
---
Apart from all of this, it's funny that you pointing to the price of implementing in ML software in C. It's not as if there are large amounts of Haskell programmers available. Writing such software in Haskell has about the same economic risk as using COBOL ;).
Still, depending on the problem C could still be the best language, no need to piss on that ...
I've found that Matlab/Octave is a decent substitute for a "high-level language" to sketch out new approaches with. They're significantly fast, as well as significantly suited to matrix algebra that they can give decent results, even though they have some less-than-beautiful code. Matlab appears to be the language of choice for AI at the University of Toronto.
Personally, I think the best option would be to roll with a functional language (or at least, a language with functions as first-class objects), since a lot of ML algorithms can be reduced to recursion on several matrices, often using very similar functions. For example, ANNs frequently have very similar structures and training strategies, but simply use different learning functions.
Everything can be done in C/C++ though, and while it'd be harder, ML is an area where the gain in speed and efficiency is so significant that extra development time pays boatloads in terms of ROI. Even basic ML examples often involve dealing with 300x236 dimensional data, so you can imagine how much that data would scale up significantly in production environments.
Don't forget machine learning generally involves a lot of experimentation, and this is easier with higher-level languages. Hand-optmization always makes assumptions about specific data structures and details which can be hard to change later.
You can get very close with code-generating high-level languages though. See Theano (http://deeplearning.net/software/theano/), for example. It uses a functional approach to compose the formulas for (c-)ANNs, and has advanced features such as automatic differentiation. Scalable generated GPU code can beat the pants of even the fastest hand-written C loops. Sure, hand-optimized GPU code can perform even faster but in my experience that is usually not worth the trouble.
The language features that the interpreter supports is subtly (and sometimes not so subtly) different from the compiled version though they share the same syntax.
In the context of scaling machine learning code you often hear that one should/could write most of it in matlab/octave and the critical parts in C. But anyone who has done it would know it is such a butt-hurting nuisance. In comparison the C integration is a pleasure in Lush.
Fortran is frequently used for physics simulation. We used Mathematica, Maple and Matlab even in non-computing classes.
I am working on a calculus refresher right now (I really need it) and am planning on working through an elementary linear algebra text once that is done, which is why I particularly noticed the comment. I have found that having an idea of applications helps retention, which is one reason I am trying to track down something more concrete.
http://www-math.mit.edu/~gs/papers/starting2matrices.pdf
It's the most accessible starter I've found to date. There is also a linear algebra group-learning thread somewhere on HN.
Then again i did not code it with performance in mind. I overloaded algo.generic for matrix operations, so the fixnum math was probably not inlined.
I've been playing around with R for the past two weeks and have been more or less happy (with the exception of the memory and speed limitations in the GNU implementation).
Nevertheless, this will be a really exciting field for the next decade or so. Amazing possibilities right now!
I think the ML class at my university--it's in "beta" right now--is using Python following similar logic.
That said, Clojure is certainly a lot faster than Python for native code and the performance gap with Java is continuing to drop. It's also worth pointing out that you should not ignore C or C++ libraries just because your are running on the JVM. It's not very hard to interface to a core C library with JNI, although be warned, there is a slight trick to doing this via clojure rather than Java.
http://research.microsoft.com/en-us/um/cambridge/projects/in...
or at this:
http://projects.csail.mit.edu/church/wiki/Church
Have a great day!