Hamiltonian Monte Carlo (2021)
bjlkeng.github.io
bjlkeng.github.io
"Monte Carlo is an extremely bad method; it should only be used when all alternative methods are worse." -- Alan Sokal
The reality of the situation is that you would only bother understanding and implementing such a complex algorithm if you had a rather difficult distribution to sample. Unfortunately, that means that you are very unlikely to know whether the answers you are getting are at all correct. You may have a bug in your code (highly likely), or the method's hyper parameters may be poorly tuned (likelier still), or the MCMC chain is not mixing rapidly leading to undetectable approximation error (almost certainly).
Almost always I would recommend attempting to simplify whatever model you are approximating and to use a method where approximation quality can be assured, rather than attempting to approximate complex (high dimensional) distributions with a more efficient algorithm like HMC.
Of course, sometimes you must approximate such models, and for this you should be glad the method exists, and good luck to you.
There was a period a few years ago when it was all the rage to take arbitrary probabilistic programming models and just toss ADVI at it blindly with fancy tools like Pymc3 or stan. I feel like everybody eventually came to the conclusion that if the model was simple enough that you could guarantee that ADVI was actually correct, you didn't need it, and if it wasn't, you couldn't possibly verify you were approximating anything close to the true posterior.
At least with HMC it would generally explode when dealing with pathological geometry (multimodality, non-identifiability, whatever), whereas lots of the approximate methods will 'converge' to completely incorrect answers. I know there's been some work on determining if things have gone off the rails (https://arxiv.org/pdf/1802.02538.pdf), but I couldn't ever find a place where it felt safe to use this stuff blindly.
Superior in what ways?
I could imagine simplicity for many problems. Roughly "here, replace your complex simulation with this magic function approximator". But for problems where you're calculating a physical quantity with a true answer, I don't see NNs being "much superior".
For large multilevel GLMs I have found Rhat and posterior predictive checks to be sufficient to diagnose HMC sampling issues.
But, of course, these might not fall within the category of difficult distributions you mentioned.
I think what anyone should be cautious about is thinking that these more sophisticated methods can allow MCMC methods to significantly scale up to larger problems. But they are probably 2-10x at best, if it's well suited to your distribution (however you characterize that). Often it's not enough to keep up with the dimensionality of the models people want to use. I remember struggling to get this to work on a Bayesian neural network that was maybe 100 parameters, tiny by todays' standards.
As for approximating hyper-parameters, again if it's an easy problem, sure, probably something will work. But I do have to say I find this really fascinating when you start to get into these harder problems with MCMC algorithms. The first thing that comes up is how do you know that a sampler is drawing samples from the correct distribution? how exactly do you quantify that? Usually each sampler takes the same amount of time to produce a sample, so what you're really comparing is which one is giving "better" samples. But.. they are all supposed to be drawing from the same distribution, so how can one sample be better than another. If they are not drawing from the same distribution, then how can you possibly know which one is the right one!
So I'm simplifying the problem here, and not to disparage any of the fine work on autocorrelation and mixing rates and all that, but you have to admire just how impossible this is generally.
I have no experience with those tuners, but, absolutely. There's no way that some heuristics can generalise to every possible distribution you feed in. But if the distribution is "sufficiently" nice (which is your responsibility to ensure) then presumably they will work.
Source code: https://github.com/chi-feng/mcmc-demo/blob/master/algorithms...
It also has visualizations for other flavors of Hamiltonian Monte Carlo, including the No-U-Turn Sampler (NUTS) used by Stan.
My perspective is that MCMC is one of the great scientific discoveries of the 20th century. There are pros and cons to different samplers, lots of caveats, concerns over convergence—all valid. Still, MCMC opens up a whole new world of model building, letting you easily try a whole slew of models. It provides flexibility, even creativity. As a concept it’s clever. As a technical achievement it’s amazing.