HNHacker News
TopNewBestAskShowJobs

preygel

50 karma · joined February 3, 2012

http://tolypreygel.com
submissionscomments
preygel··on Julia 0.5 Highlights
Typically a mix of reasons:

- Needing to interact with an existing codebase, and an existing developer base. If everyone knows and uses Python and only a few use Julia, it is too early to put Julia in production. If there are proprietary libraries, now may not be the best time to commit to porting them to Julia.

- Language and ecosystem stability. I started something on 0.4, and with 0.5 there were a raft of deprecations. If the code will live several years, that's a support commitment with unclear value.

- Library maturity. If I need to build a web app, read an Excel, read a CSV with dates quickly, consume a SOAP endpoint, etc etc in Python -- no problem. With Julia I will mostly be fine, but am likely to run into some cases that are not yet 100% there.

- Most code does not need the extra performance, so once you have a fast prototype as a performance target it is often not that hard to hit similar performance with Python + numba/Cython.

Note for that last point: there is a lot of value to not worrying about this in the exploratory stage, and getting a performance target (for later optimization) as a nice byproduct.

preygel··on Julia 0.5 Highlights
Exciting to hear -- I have used Julia for prototyping, and have found it to be excellent for that:

In my experience there are still some rough edges as compared to the Python ecosystem (of course!), which together with the 0.x status make it impractical for many production situations. However it is fantastic for prototyping numerical code, the type system is a pleasure, and the JuMP mathematical optimization library is a gem. Being able to have fast code be "first class," as opposed to the impedance mismatch of dealing with numba/Cython, feels great and is a real boon for trying new things. Then, it is fairly straightforward to port the final solution to whatever production language you use (e.g., Python with a sprinkle of numba).

preygel··on Image unshredding using a TSP solver
Absolutely!

Not to mention the often impressive performance of "general" MIP solvers. It is only a shame that the best ones there are commercial (Gurobi followed by Cplex). That said, Cbc is lovely in a wide range of cases, and is open-source.

preygel··on Image unshredding using a TSP solver
A bit tangentially, this is also a great display of just how good readily available approximate solvers have gotten for a wide range of combinatorial optimization problems (like TSP).

The LKH solver used here has rather impressive performance. From http://webhotel4.ruc.dk/~keld/research/LKH/: "LKH has produced optimal solutions for all solved problems we have been able to obtain," and the studies linked there show that it really does find optimal solutions "with an impressively high frequency."

While the point here is to use an off-the-shelf solver, it can be nice to have visibility into what's actually happening:

- This is a reasonable example for explaining some of the local search TSP heuristics. For instance a "2-opt" move corresponds to picking a contiguous range of columns and flipping them.

- The previous HN post used simulated annealing, which would not be an outright terrible approach to TSP itself -- were it not for the better Lin-Kernighan-based approaches (like LKH).

- If we want, we can tell LKH to start from Sangaline's "nearest-neighbor" approach ("INITIAL_TOUR_ALGORITHM = NEAREST-NEIGHBOR"). This does not make a difference here though.

preygel··on How to Become a Data Scientist, Part 2
The distinction is being made on end outputs -- is it more like a Powerpoint deck, or more like a production data pipeline -- not on techniques.

"Machine learning" includes plenty of activities that can be used to provide evidence for one-off human decision making (e.g., using a model to produce forecasts or to understand sensitivities).

preygel··on Habits of highly mathematical people
Math PhD here. I have also had the chance to work with many bright "analytically-minded" people from other backgrounds while a management consultant.

This article is spot on, and some of the behaviors really do seem more indicative of "mathematical people" -- which I suggest really stands for "those who have done research in a 'mathematical' field." (The key being the mix of cold, hard precision in the idealized proof with the squishy, intuitive, human activity of discovering what is pretty and true -- an aspect usually lost in math education!)

Some examples I found most poignant:

- The article mentions "fluidity with definitions" and illustrates it well with the anecdote about Keith Devlin. This is a skill distinct from pure "analytical reasoning," as it requires comfort with definitions that are at once precise but also open to (frequent) change. The process of forming and changing definitions is creative and imprecise, and falls into what is sometimes called "conceptual reasoning." (A programming analog might be API design.)

- Several of the other points are tools for figuring out what is true, and for precising imprecise statements. For example the need to "teas[e] apart .. assumptions" is only natural when reading papers with Theorems that have very precise conditions .. which do not exactly hold in the case you need! In many other "analytical" contexts pre-conditions are not made as precise and arguments by analogy are considered acceptable provided the conclusion is believed. (A programming analog might be debugging when some implicit pre-conditions or invariants break.)