1,357 karma · joined March 28, 2007
I think writing some algorithms and using them to solve problems provides great motivation for the math. In particular, the math will explain why certain approaches did and did not work. Without the hacking that material can get a bit dry.
Beyond that we are seeing some cool stuff with Blackbox Variational Inference [3] and other Automatic Variational [4] solvers coming out of Blei's lab. At present these capabilities are spread about the different system but I expect all them to eventually end up available in whichever you choose to use.
[1] http://homes.soic.indiana.edu/ccshan/rational/simplify-padl....
[2] http://www.srl.inf.ethz.ch/papers/psi-solver.pdf
[3] http://www.cs.columbia.edu/~blei/papers/RanganathGerrishBlei...
[1] https://web.stanford.edu/~ngoodman/papers/aistats2014-shred....
[2] http://docs.webppl.org/en/master/inference/methods.html#enum...
I think Figaro, Stan and PyMC are the most "production-ready" in the sense they have been used for projects outside the realms of their creators. Still I would argue on some level all of them are research projects that aim to explore how to make probabilistic modeling more accessible to people. Ideas in one language often will appear in another down the line. So I encourage to explore a few of them and reach out to the people working on them.
Good luck and I hope the work is published.
http://www.jetpens.com/Zebra-Mackee-Wet-Erase-Double-Sided-M...
http://journals.aps.org/prl/pdf/10.1103/PhysRevLett.116.0611...
https://dcc.ligo.org/public/0122/P1500217/014/LIGO-P1500217_...
https://software.intel.com/en-us/blogs/2016/02/14/python-bri...
In some sense, python and ruby compile into bytecode, and then the bytecode is executed using an interpreter. There are also concepts such as partial evaluation (where you do a little interpreting while you compile) and JITing (where you do a little compiling while you interpret) which also seem to mix compilation and interpretation, but can be clearly specified as a particular combination of them.
MCMC is a very slow inference algorithm. Its primary advantage was that for well-known models it could be coded up much more simply than a fancier inference technique. When you consider variational methods and newer streaming methods based on things like Assumed Density Filtering you can get really great scalable performance. The point of probabilistic programming is write inference algorithms once for a large class of models and be done. So the advantage of using a fancier method is amplified.
This means paradoxically probabilistic programming should eventually be faster than existing methods rather than slower, since you can reuse these fancier inference methods for new models. This is a very active field so this progress is only starting to be appear in the existing systems.
As an example, suppose I am trying to load a csv file with two fields being strings and one being Double. A list of lists representation isn't going to cut it. Instead I should make a Row datatype
data Row = Row {field1 :: String, field2 :: String, field3 :: Double}
Then you can parse the csv file into a [Row] representation.As an aside, if the task did involve parsing csv I suggest the cassava library which I found out about through the amazing What I Wish I Knew When Learning Haskell (http://dev.stephendiehl.com/hask/)