HNHacker News
TopNewBestAskShowJobs

idunning

826 karma · joined March 26, 2012

http://iaindunning.com
submissionscomments
idunning··on GitHub and Jupyter IPython Notebooks
It is an interesting point. In my usage, I don't view it as a one-or-the-other proposition. I use Julia, so if I'm doing some exploratory work and experimentation, or making something to present results, then the IJulia notebook is great. If I'm writing some serious longer-running stuff, or a package, I'm in an editor. Sometimes I'll write code in a separate file and call it from a notebook just to keep the notebook focussed on communicating something, and "hiding" the details.
idunning··on The Traveling Salesperson Problem
Great stuff from Norvig, as per usual. I find it strange that the next step after enumeration is a heuristic approach, something that comes up a lot in "solve a TSP" posts - there are many ways to solve TSP to provable optimality that perform really well on non-pathological instances, and can often be stopped at any point to get both a probably near optimal solution and, even better, a bound on how good the best solution could be. I would argue that the implementation cost is outweighed by the possibility of usually getting optimal solutions and by having confidence in how far off you are (no chance of producing a terrible solution).
idunning··on Route Optimizer written in Julia
The really unfortunate thing about that post is it completely fails to understand the essential part of solving the TSP this way, which are the lazily-added subtour elimination constraints. I don't think the blog author at Forio really understood the code they were using, hopefully they will update it one day!
idunning··on Route Optimizer written in Julia
I wrote the solver code used here as an example for the JuMP modeling package for Julia, which is available at https://github.com/JuliaOpt/JuMP.jl. The specific TSP example code is available at https://github.com/JuliaOpt/JuMP.jl/blob/master/examples/tsp..., and more general information about optimization in Julia is available at http://juliaopt.org
idunning··on Reverse Engineering Wipeout
Amazing, I've tried to do this for Wipeout 3 in the past (extract the tracks) and was defeated - this makes it seem fairly easy!
idunning··on Announcing Rust 1.0 Beta
CI against package ecosystems is a really great idea. We do it for Julia too [1], and it can identify some really subtle issues that would otherwise take longer to become apparent.

http://pkg.julialang.org/pulse.html

idunning··on Deep Learning, the Curse of Dimensionality, and Autoencoders
Original content is at http://nikhilbuduma.com/2015/03/10/the-curse-of-dimensionali...
idunning··on Beginning deep learning with 500 lines of Julia
I don't think its the same, but there is a high-quality implementation in Julia that is registered: https://github.com/dfdx/Boltzmann.jl
idunning··on Machine Learning Done Wrong
Some techniques, e.g. random forest, give variable importance indicators for free. If you can test it out, give it a go - don't have to use the random forest as the final model.
idunning··on Machine Learning Done Wrong
If you are disciplined, and separate data into training and testing sets, you can try as many models as you want without fear of overfitting. Indeed, optimizing over the parameters of a model on the training set is essential (pruning parameters in a tree, regularization weights, etc.) and can be thought of as training large number of models.

If you aren't doing this correctly, then you can't really interpret the performance of even a single model. Seen people screw this up in so many ways - my favorite recent one that was quite high on HN was someone using the full dataset for variable selection, before doing a training-testing split afterwards.

idunning··on Beginning deep learning with 500 lines of Julia
The author mentions that one of his goals was to focus on a smaller set of functionality and make it simple and high-performance, but I've got to put a shoutout to Mocha.jl here [1]. It is essentially Julia's answer to the Caffe deep learning framework (which is linked in the article), and has pure Julia, C++, and CUDA GPU backends. Its under active development but is already pretty amazing. Bonus: it has documentation!

On the contents of this blog post: I really like how the Julia type system is used here. Not only do the types help structure the code and send a signal to the user, but of course there is type-checking to catch errors.

[1]: https://github.com/pluskid/Mocha.jl

idunning··on OPL – High-level syntax for linear programming
PuLP, apart from being in Python and this being in Ruby, is definitely a more general tool as it can connect to a wide variety of (MI)LP solvers.

Heres a link to the latest version of PuLP: https://pypi.python.org/pypi/PuLP/1.5.6

idunning··on OPL – High-level syntax for linear programming
An alternative is JuMP [1], a package for Julia. You can model optimization problems with linear, quadratic, and general nonlinear constraints/objective and send them to a variety of open-source and commercial solvers. JuMP, these solver wrappers, and more are all part of the JuliaOpt organization [2]. Once of the great things about JuMP is that it generates the internal/computer-friendly form for LPs/MILPs/MIQCQPs/... very quickly, as good as any commercial tool. If you are solving nonlinear problems, it'll generate first and second derivatives for you using automatic differentiation (see JuliaDiff for more on that [3])

[1]: https://github.com/JuliaOpt/JuMP.jl

[2]: http://juliaopt.org

[3]: http://juliadiff.org

idunning··on Quantitative Economics with Julia [pdf]
In my opinion, as a contributor to Julia and someone who teaches machine learning with R - start with R. Things will "just work" for the most part and you won't have to worry about whether your packages will work while you are learning ML. I recommend using the "caret" package in particular: it puts all the ML packages behind a nice common interface and has goodies like crossvalidation and train/test splits built in.

Python with Scikit-learn could be a good choice too from everything I hear (possibly even better, by some accounts).

To be clear, Julia is more than capable of doing ML, but I'd say that interface-wise its not quite there yet. Most of the pieces are there, everything from DataFrames to wrappers for GLMNet to random forests, and even the deep learning library Mocha.jl (check it out, its fantastic!). If you were to implement a new ML algorithm, I'd want to be doing it in Julia - it'll perform great without having to get in a multi-language scenario (like R+Rcpp or Python+???[numba?]).

idunning··on A simple Minecraft written in Rust
How is that dependency graph plotted in the README?
idunning··on Julia is awesome, but..
Are you talking about compiling R from source, including all dependencies and high-performance BLAS, LAPACK, FFT, ARPACK, etc. libraries? If not, you are comparing apples to oranges.
idunning··on Julia is awesome, but..
I feel like people are perhaps a bit overly negative about the state of Julia's packages. Compared to other early-stage languages, I'd say our package ecosystem is vibrant and full of fantastic packages. Naturally perhaps they are more math/scientific focussed, so if you are looking for cutting-edge packages for web dev you won't find them yet (although there are packages!).

See http://pkg.julialang.org/pulse.html, for example. We have over 470 packages in total that are registered, and on Julia 0.3 we have over 300 packages with tests that pass - and we run the tests in all registered packages every night.

Some of my favorite packages (that I didn't make, of course :D) would include

https://github.com/JuliaStats/Distributions.jl

https://github.com/JuliaStats/StatsBase.jl

https://github.com/pluskid/Mocha.jl (deep learning)

https://github.com/stevengj/PyCall.jl

The JuliaOpt stack of optimization packages (http://juliaopt.org)

and then you get fun new ones like https://github.com/anthonyclays/RomanNumerals.jl

idunning··on Should I Get a Ph.D.?
Countering your counter-point: this seems like the most pessimistic comment about doing a PhD I've ever seen. There are many types of department, many fields, that don't operate like the type of setup you are describing. I'd be damn surprised if anyone in my department is waking up with "screaming nightmares" or is "on their way to the psychiatric hospital" - though maybe I'm just ignorant about my peers, or got a lucky roll of the dice.

I'm not saying this doesn't happen, or trying to diminish your personal experience, but you've presented your dark scenario as being as inevitable as the happy scenario you are railing against.

idunning··on Julia Package Ecosystem Pulse
I should point out the main Julia package listing is at http://pkg.julialang.org/, this is the "executive summary" page. At the bottom of this "Pulse" page are test results - we run the tests for every package every night on both Julia stable (0.3) and unstable (0.4) to make sure everything is still working nicely together. Although, breakages on 0.4 just mean cool stuff is happening on master.
idunning··on Visualising dependencies in Go
I tried something similar with Julia packages [1] but I only tried force graphs - with similarly poor results for the most part. The data I used is already out of date (got another 80 packages since then! [2]) so maybe a good reason to revisit with these other graph styles.

[1]: http://iaindunning.com/2014/pkg-deps.html

[2]: http://pkg.julialang.org/pulse.html

idunning··on Beyond Light Table
I'm a PhD student, and a big Julia fan (and contributor), and I STILL think everyone should know Python (or similar). It truly can be the language for everyone, and I've found it amazingly easy to teach it to people.
idunning··on Using Machine Learning and Node.js to detect the gender of Instagram Users
Neural networks have their place, but are probably the most complicated and opaque machine learning tool. They are also hard to set up: so many parameters! Given that, I found it really strange that they went straight for a neural network (and then implemented one themselves!). Surely the place to start would be Naive Bayes, followed up with regularized logistic regression (through something like glmnet). Heck, even random forests would do quite well on this task I imagine, although thats getting closer to on the complexity and opaqueness spectrum towards NN.

There is also no evidence of doing cross-validation, and in another comment they say they used entire data set to do variable selection - a pretty bad mistake. They justify by saying they aren't in an academic environment, but thats kind of a bad excuse, as given the way they've done it I'm very unsure whether they actually are getting the accuracy they think they are.

I also worry that they sunk two man-months into this when they could probably have achieved similar if not better results with off-the-shelf and battled-tested tools. Sets off a lot of warning bells.

idunning··on Complex Step Differentiation
Ah that makes a lot of sense!
idunning··on Automasymbolic Differentiation
There is a lot of support and interest in automatic differentiation [1] in the Julia [2] community, partly because the language design makes it relatively easy to do. In fact, there is a whole "organization" dedicated to AD packages, JuliaDiff [3]. In particular there are packages for dual numbers and their generalizations, as well as reverse-mode AD packages. ReverseDiffSparse.jl, for example, uses some clever tricks including graph coloring to create very efficient Hessian matrices.

You can make use of AD for more than just playing around too, esp. for optimization (JuliaOpt [4]): Optim.jl will use them to calculate exact derivatives if you don't provide them, and JuMP.jl will use them to calculate the sparse Jacobian and Hessian matrix for a nonlinearly constrained optimization problem (which can be used by, e.g. Ipopt.jl)

[1]: http://en.wikipedia.org/wiki/Automatic_differentiation

[2]: http://julialang.org/

[3]: http://juliadiff.org/

[4]: http://juliaopt.org/

idunning··on Complex Step Differentiation
This is very similar to dual numbers [1] which I find even easier to reason about.

There is an implementation of dual numbers in Julia [2] that is quite fun to play around with. The Optim.jl package [3] uses this to get better derivatives than finite differencing.

[1] http://en.wikipedia.org/wiki/Dual_number

[2] https://github.com/JuliaDiff/DualNumbers.jl

[3] https://github.com/JuliaOpt/Optim.jl

idunning··on Tales of Statisticians: George B. Dantzig
Simplex method was a game changer.

Its a funny algorithm: worst-case time complexity is exponential in the input size, and is pretty easy to demonstrate (see "Klee-Minty cube" on Wikipedia). It has a couple of "rules" you can change out that will fix the exponential problem for some cases, only to introduce new problematic cases elsewhere. However the reality is that it demonstrates polynomial-like performance on almost all problems of interest, and that is why it is so widely used.

Later, interior point methods arose that can be faster sometimes (and are polynomial time complexity), but they didn't kill the simplex method. This is partly due to one key property of the simplex method: at optimality, you can change the linear program in many different small ways and start the algorithm again from where you left off (sometimes you need the dual simplex method, a sibling method). You'll return to feasible optimality in usually only a few iterations. This is what powers the branch-and-bound approach to integer programming, which is the really useful application of LP these days. Interior point methods don't really have good warm starts to this day, certainly not good enough for branch-and-bound.

I've met several professors who kinda don't like the simplex method because (they say) it is not a beautiful algorithm from a theory perspective, but I think its wonderful.

Oh, and PSA: very difficult to implement correctly! The textbook algorithm will fail terribly on real problems due to floating point issues - please use an existing implementation if you need to solve LPs!

idunning··on Juliabox
Thanks so much! I was going to flag but was unsure if it was appropriate.
idunning··on Juliabox
The site wasn't meant to be publicly released yet - it was online only for internal testing amongst a small set of users. Its under very heavy load now so is most likely unresponsive. You can get a similar experience by trying out IJulia on your own machine though, until its properly released.
idunning··on Juliabox
Yes it was posted by a third-party, not fully ready for primetime. Its running on AWS so at least no one's computer is likely to get owned due to any glitches! Thanks for your interest - it is a pretty fun tool.
idunning··on Juliabox
Yes indeed, the next evolution of the concept. As you say, the IJulia notebook is the "right" way to do it.
← PreviousPage 2 of 5Next →