826 karma · joined March 26, 2012
If you aren't doing this correctly, then you can't really interpret the performance of even a single model. Seen people screw this up in so many ways - my favorite recent one that was quite high on HN was someone using the full dataset for variable selection, before doing a training-testing split afterwards.
On the contents of this blog post: I really like how the Julia type system is used here. Not only do the types help structure the code and send a signal to the user, but of course there is type-checking to catch errors.
Heres a link to the latest version of PuLP: https://pypi.python.org/pypi/PuLP/1.5.6
[1]: https://github.com/JuliaOpt/JuMP.jl
[2]: http://juliaopt.org
[3]: http://juliadiff.org
Python with Scikit-learn could be a good choice too from everything I hear (possibly even better, by some accounts).
To be clear, Julia is more than capable of doing ML, but I'd say that interface-wise its not quite there yet. Most of the pieces are there, everything from DataFrames to wrappers for GLMNet to random forests, and even the deep learning library Mocha.jl (check it out, its fantastic!). If you were to implement a new ML algorithm, I'd want to be doing it in Julia - it'll perform great without having to get in a multi-language scenario (like R+Rcpp or Python+???[numba?]).
See http://pkg.julialang.org/pulse.html, for example. We have over 470 packages in total that are registered, and on Julia 0.3 we have over 300 packages with tests that pass - and we run the tests in all registered packages every night.
Some of my favorite packages (that I didn't make, of course :D) would include
https://github.com/JuliaStats/Distributions.jl
https://github.com/JuliaStats/StatsBase.jl
https://github.com/pluskid/Mocha.jl (deep learning)
https://github.com/stevengj/PyCall.jl
The JuliaOpt stack of optimization packages (http://juliaopt.org)
and then you get fun new ones like https://github.com/anthonyclays/RomanNumerals.jl
I'm not saying this doesn't happen, or trying to diminish your personal experience, but you've presented your dark scenario as being as inevitable as the happy scenario you are railing against.
There is also no evidence of doing cross-validation, and in another comment they say they used entire data set to do variable selection - a pretty bad mistake. They justify by saying they aren't in an academic environment, but thats kind of a bad excuse, as given the way they've done it I'm very unsure whether they actually are getting the accuracy they think they are.
I also worry that they sunk two man-months into this when they could probably have achieved similar if not better results with off-the-shelf and battled-tested tools. Sets off a lot of warning bells.
You can make use of AD for more than just playing around too, esp. for optimization (JuliaOpt [4]): Optim.jl will use them to calculate exact derivatives if you don't provide them, and JuMP.jl will use them to calculate the sparse Jacobian and Hessian matrix for a nonlinearly constrained optimization problem (which can be used by, e.g. Ipopt.jl)
[1]: http://en.wikipedia.org/wiki/Automatic_differentiation
[4]: http://juliaopt.org/
There is an implementation of dual numbers in Julia [2] that is quite fun to play around with. The Optim.jl package [3] uses this to get better derivatives than finite differencing.
[1] http://en.wikipedia.org/wiki/Dual_number
Its a funny algorithm: worst-case time complexity is exponential in the input size, and is pretty easy to demonstrate (see "Klee-Minty cube" on Wikipedia). It has a couple of "rules" you can change out that will fix the exponential problem for some cases, only to introduce new problematic cases elsewhere. However the reality is that it demonstrates polynomial-like performance on almost all problems of interest, and that is why it is so widely used.
Later, interior point methods arose that can be faster sometimes (and are polynomial time complexity), but they didn't kill the simplex method. This is partly due to one key property of the simplex method: at optimality, you can change the linear program in many different small ways and start the algorithm again from where you left off (sometimes you need the dual simplex method, a sibling method). You'll return to feasible optimality in usually only a few iterations. This is what powers the branch-and-bound approach to integer programming, which is the really useful application of LP these days. Interior point methods don't really have good warm starts to this day, certainly not good enough for branch-and-bound.
I've met several professors who kinda don't like the simplex method because (they say) it is not a beautiful algorithm from a theory perspective, but I think its wonderful.
Oh, and PSA: very difficult to implement correctly! The textbook algorithm will fail terribly on real problems due to floating point issues - please use an existing implementation if you need to solve LPs!