An Introduction to Probabilistic Programming
arxiv.org
arxiv.org
[1] http://web.stanford.edu/~ngoodman/papers/POPL2013-abstract.p...
That said, you can get pretty far with probabilistic programming in any language with decent monad support [2]. I did most of my probabilistic programming work in Scala. (You lose the ability to do really fancy inference if you go the monad route, as you can't analyze the program structure, but a lot of the time, this is fine.)
Would make it much easier to play around with the language when trying to wrap my mind around the church version of probmods.
I’ll have a look on google on the names you mentioned :)
^{:nextjournal/viewer :plotly}
{:data [{:x '(0 10 20 30 40)}
{:y '(5 10 15 20 25)}]
:layout {:autosize false :width 800 :height 400
:xaxis1 {:title "year"}
:yaxis1 {:title "revenue"}}}
See the result here:https://nextjournal.com/kommen/plotting-from-clojure
Regarding anglican, here is a link. It’s basically a church port:
on that topic, can anyone recommend an online stats/probability course? I tried the coursera one by Sebastian Thrun and couldn't get far into it because the "TA" examples were unintelligible.
https://www.youtube.com/playlist?list=PL05umP7R6ij1tHaOFY96m...
[1] https://dtai.cs.kuleuven.be/problog/
[2] https://github.com/ML-KULeuven/problog
[3] https://github.com/ML-KULeuven/deepproblog
This gives a nice picture of what's happening. At the same time, does this mean that in the end, you're basically operating on a single distribution with only a few canned global transformations?
On the surface probability seems antithetical to the explicit well-defined determinism of programming.
But if you modify the question to ask how likely it is for 3 people, or you add things like February 29th and leap years, or you add the fact that births are more likely during summer months, then it becomes extremely difficult to solve this deterministically. Instead, you run a Monte Carlo simulation to get approximate probabilities. This is much simpler to code and can be easily modified to fit new conditions.
A Monte Carlo simulation is probabilistic because you use random numbers in the simulation and you won't get a 100% perfect answer but you'll get close enough (and you can do some math to get error bounds).
Determinism is a nice illusion that quickly breaks down on real data. Scientific problems are non-unique and noise or inadequate models cause data and model to not mesh with one another. Deterministic answers have the property of being precisely wrong as opposed to mostly correct.
An HTTP library is good for making web calls. A probabilistic program is good for estimating probabilities.
One example would be polling. You could write down a program that estimates votes from political polls. Then feed it polling data and get estimates of how people will vote.
Also, randomized algorithms for all kinds of things are very useful in general-purpose computing
I remember it being pretty hyped 5-6 years ago …
PPLs using MCMC (e.g., Stan and pymc) are first-choice tools for Bayesian inference on small to medium data analysis, which is lots more common than Google-scale data analysis.
"Every little kid knows that even the slightest variation in the placement of a firecracker or the most seemingly minor imperfection of a glue joint will lead to dramatically different model airplane explosions."
I've never made my model airplane explode (after the many hours needed to build them).
But I burnt ants :-)
> It is a Lisp-like language which, by virtue of its syntactic simplicity, also makes for efficient and easy meta-programming, an approach many implementors will take. That said, the real substance of this book is language agnostic and the main points should be understood in this light.
You can easily do this calculation by hand or in Python, but this does not generalize to more complex real-world scenarios. For complex probabilistic models, we must rely on numerical approximations. MCMC is just one algorithm for doing this approximate inference. Another popular technique is called variational inference [2]. Another commenter mentioned HMC [3], which is just a specific instance of MCMC.
Basically probabilistic programming is a way of describing a distribution, and then MCMC is one way of inferring the quantities in that distribution.
If you're taking a Bayesian approach to statistical modelling and inference, then they're probably a fairly good tool to consider. With the Bayesian approach you're trying to compute some posterior probability distribution that summarises your prior information (this might capture domain knowledge, information from related studies) and information from observations.
There are different ways to compute a posterior distribution. In very simple or contrived cases you might be able to manually grind out an answer analytically with pen and paper and lots of algebra and integrals. But that isn't very efficient or scalable. Also, nice algebraic structure is very easily broken by small perturbations to the problem statement -- need to add a weird bit onto the model to capture some real world behaviour? Good chance that ruins your algebraic structure and previous analytic "attack".
MCMC can be used to estimate the integrals you need when computing a posterior distribution. MCMC isn't the only way to estimate or approximate these calculations -- e.g. another approach is variational inference where a bunch of approximations are introduced to replace the original calculation with an approximation that is easier to compute -- this likely introduces bias into the results but can give you something that can then be solved analytically or semi analytically (e.g. approximate everything as Gaussian distributions and a lot of integration collapses to efficiently computable algebraic identities).
Some probabilistic programming platforms like Stan let you define your probabilistic model and parameters and decouple it from the computational backend used to estimate the posterior distribution. E.g. in Stan you can switch the computational backend between MCMC (https://mc-stan.org/docs/2_18/stan-users-guide/sampling-diff...) and ADVI (auto-differentiation variational inference).
MCMC has practical problems in that it is only guaranteed to give you the correct (unbiased) estimate asymptotically, in the limit if you run it for an infinite amount of time. If you're trying to approximate the integral of a function that is very multi-modal -- where it would be difficult for a global optimisation algorithm to locate the global optima -- then MCMC will likely also struggle to produce a good estimate. MCMC is difficult to parallelise effectively as the algorithm is inherently like an iterative local search procedure -- the next state in the chain is some mutation of the previous state. You can run n MCMC chains in parallel from n different initial configurations, but it's not obvious that you'll get a better estimate from n short chains vs a single long chain -- the longer a chain runs, the more chance it has of being able to discover and explore higher probability (more realistic, more plausible) configurations of the parameter space.
MCMC isn't only used for probabilistic programming, you can apply it for other things. E.g. it gets used in material science to study statistical properties of molecular dynamics simulations etc.