GoLearn – Machine Learning in Go
sjwhitworth.com
sjwhitworth.com
Numpy and Scipy are so mature in this respect it is difficult to compete with them. I looked into implementing eigenvalue algorithms recently with the idea that I would just write a native Go library for doing this kinda stuff. However, reading the source of JAMA[1] was sufficiently humbling for me to realize this was not a good idea. (If you really want to be humbled try reading the fortran implementations in LAPACK.[2] I believe SRC/dgeesx.f is a good starting point)
[1] http://math.nist.gov/javanumerics/jama/ [2] http://www.netlib.org/lapack/#_software
Tell me about it. What I'd give to have something equivalent for Objective-C (or certain other languages too, e.g. Julia). I'm looking at PyObjC as a stop-gap solution for now, but it sure adds complexity to a project.
Edit: Since SciPy is BSD-licensed and presumably mostly C behind-the-scenes, perhaps there's potential for a group to try and package it up for other languages? I have no idea how large an undertaking like that would be...
I don't know how mature they are though.
Thanks for the pointers!
EDIT: The lack of README and documentation beyond the API docs concerns me for mat64. still a pretty interesting project. Might be useful for non-critical stuff.
Even just KNN needs several distance metrics built in (manhattan, hamming, mahalonobis, to name a few) and a good ball tree implementation for use on large datasets so that the search time goes from N to log N.
There are already a couple of Machine Learning libraries[1][2] written in Go and some of them are actually more mature than GoLearn.
Also just curious, I always thought Go is not really a good language for DM/ML stuff due to lack of good matrix library and generics. If someone here actually tried to write any ML library in Go, what's your genuine feeling about it?
[1] https://github.com/huichen/mlf [2] https://github.com/xlvector/hector
It's nice being able to trivially parallelise operations in Go - e.g. constructing the weak learners for a random forest, generating candidate splits, recursing down left and right branches, etc.
// Recur down the left and right branches in parallel
w := sync.WaitGroup{}
recur := func(child **pb.TreeNode, e Examples) {
w.Add(1)
go func() {
*child = c.generateTree(e, currentLevel+1)
w.Done()
}()
}
recur(&tree.Left, examples[bestSplit.index:])
recur(&tree.Right, examples[:bestSplit.index])
w.Wait()
As you said, generics and a matrix library would be make the experience nicer. Just having sort :: Ord a => [a] -> [a]
would strip a decent amount of mildly error-prone boilerplate, and there are other cases (splits for cross-validation, etc) where it would be nice to be able to abstract over the type of the slice, etc.The language and tooling (pprof, go fmt, go doc) are great and make it quick to write and optimize stuff so it is well suited for my (largely experimental) purposes.
I also really like slices for writing efficient code as they let you pre optimize and reuse arrays and not have to keep track of the ending position.
Matrix libraries would be nice but you can call c ones via cgo. I am hopping for efficient pure go ones to be developed eventually so you can use them on app engine/nacl/exacycle or other untrusted code environments.
The comparison doesn't really work anyway - better would be Julia vs. Go, some ML toolkit in Julia vs. this.