I'm slightly disappointed, though. Almost all of it is dedicated to introducing prerequisite terminology (prior, posterior, likelihood, markov chains) which probably a lot of readers will already be familiar with, and then the actual explanation of MCMC is just this:
> To begin, MCMC methods pick a random parameter value to consider. The simulation will continue to generate random values (this is the Monte Carlo part), but subject to some rule for determining what makes a good parameter value. The trick is that, for a pair of parameter values, it is possible to compute which is a better parameter value, by computing how likely each value is to explain the data, given our prior beliefs. If a randomly generated parameter value is better than the last one, it is added to the chain of parameter values with a certain probability determined by how much better it is (this is the Markov chain part).
"subject to some rule", "it is possible to"... I feel like this is really glossing over the actual explanation of how this works. Actually, I still have no idea why there is a Markov chain. What is the structure of the chain? Where does it come from and why can't we just sample the parameter without using a chain?
Anyway, I appreciated most of the article and it started out really promising. It would be great if the author could try to expand on that still-baffling part ;)