A Zero-Math Introduction to Markov Chain Monte Carlo Methods
medium.com
medium.com
I'm slightly disappointed, though. Almost all of it is dedicated to introducing prerequisite terminology (prior, posterior, likelihood, markov chains) which probably a lot of readers will already be familiar with, and then the actual explanation of MCMC is just this:
> To begin, MCMC methods pick a random parameter value to consider. The simulation will continue to generate random values (this is the Monte Carlo part), but subject to some rule for determining what makes a good parameter value. The trick is that, for a pair of parameter values, it is possible to compute which is a better parameter value, by computing how likely each value is to explain the data, given our prior beliefs. If a randomly generated parameter value is better than the last one, it is added to the chain of parameter values with a certain probability determined by how much better it is (this is the Markov chain part).
"subject to some rule", "it is possible to"... I feel like this is really glossing over the actual explanation of how this works. Actually, I still have no idea why there is a Markov chain. What is the structure of the chain? Where does it come from and why can't we just sample the parameter without using a chain?
Anyway, I appreciated most of the article and it started out really promising. It would be great if the author could try to expand on that still-baffling part ;)
The theory of Markov chains also provides a useful theorem, which is that (under suitable conditions) the distribution of each sample converges to a particular 'limit' distribution. In fact it converges to the so called 'stationary' distribution. It's called stationary because if you start the Markov Chain at a sample from the stationary distribution then all following samples of the chain will have the exact same distribution.
For the finite case, if you put the probabilities of transitioning from state i to state j in a matrix M then the stationary distribution is a vector p such that Mp = p (or pM = p depending on what convention is used).
The MCMC methods consist of trying to design a Markov chain which has a useful stationary distribution. This allows you to sample the stationary distribution by sampling the Markov chain.
https://jeremykun.com/2015/04/06/markov-chain-monte-carlo-wi...
The curse of dimensionality is a cause of Monte Carlo methods to be used instead of other numerical methods in high dimensional problems.
You should communicate to your audience that the advances in stochastic sampling and its application to science has revolutionize the field not because a breathrough in maths but because great key insights that allow us to explore multidimensional problems effectively with the help of computers.
Another important factor is that MCMC is used extensively in Bayesian statistics, but to explain that main application of MCMC in your blog it is necessary to introduce concepts like Bayer's theorem, likelihood, posterior density function, and prior probabilities.
There are some blogs in which all those concepts are illustrated and code in R or python is provided. Perhaps MCMC requires more than one post to be adequately explained.
To sum up, I am a follower of your posts and I appreciate the effort you take to show the main ingredients and the python code. My only desire is for you to continue improving the content of your blog for us to follow and enjoy.
Merry Christmas to you and your family from a follower in Spain.
You only need to be able to compute the relative values of the function you want to integrate (in this case a probability density) at the current point and the proposed destination. Where is this function coming from is not relevant to understand MCMC methods, they can be applied in many problems unrelated to Bayesian statistics.
Here is another (non-zero-math) introduction to the topic: https://arxiv.org/pdf/cond-mat/9612186.pdf https://www.coursera.org/learn/statistical-mechanics/lecture...
If the parameter can be sampled from the posterior directly, then no chain is usually needed, just sample it. However, most often you can not directly sample from the posterior. Note that this is not necessarily because "posterior is hard to analytically compute". Most times you know the analytical expression for the posterior distribution (up to the normalization constant) but this does not mean you can generate samples from it easily.
So, instead, you construct a Markov chain, from which you can sample easily, and which has the property that in the limit its distribution coincides with your posterior. Then you sample the chain for a long time an hope for the best.
Any chain for which you can show that its stationary distribution is the posterior works in principle. But some chains are better than others. There are also some generic ways to construct such chains, most famous are the Gibbs and the Metropolis-Hastings samplers.
Finally, do not sample blogs by "wikipedia researchers". Sample books like "Bayesian Data Analysis" by Gelman et al. They will not lie to you.
And that's why arguing about word meanings is not a good idea. I think referring to written math as just math is reasonable, and referring to it as equations also is.
Edit: grammar
https://www.goodreads.com/book/show/10672848-the-theory-that...
Articles such as this one are a great resource to help what would be a mindless regurgitator actually understand what the whole point of an MCMC method is for. The 'why' to the 'what' essentially. And you don't need math to explain that.
There's an argument that you need math to explain it because it is math -- just not what we're taught to think of as "math".
One of my gripes with the current education system is that it makes it hard to recognize when math appears in situations not explicitly involving numbers, equations, and matrices.
Sounds like circular reasoning. There's an argument you need logic, because math is logic.
Math is confused with arithmetic, which is one tiny subset of math, and notation, which isn't math at all but is merely a tool mathematicians use to communicate more clearly because some concepts are too bulky when you attempt to express them using words. The essence of math is careful reasoning, following step-by-step processes to derive new truths from existing ones. If you do that, you're doing math, and if you do it without arithmetic or notation, people will say you're doing math without math.
> There's an argument that you need math to explain it because it is math
* More specifically, it's a mathematical idea, so, in order to explain it, it would be impossible not to use other mathematical ideas. These other ideas exist throughout the article and are [incorrectly] separated from what the average person thinks when they hear "math" but it's math, nonetheless. For instance, the many graphs throughout the article should be recognized as math.
So I guess that is not only in math, but in the entire corpus of knowledge, our system makes us prepared for grading tests, not for applying the knowledge.
As for "intuition", I usually seek refuge in Jeff Raskin's theorem "intuitive = familiar" - i.e. you don't lose the content of a statement made using the word intuition if you replaced it with the word familiar .. but now suddenly it is less "mystical" and clearer. Plus it points. To the subjectivity of intuition. - that what is familiar to one may not be familiar to another and so there are different "intuitive" explanations.
So "intuitive understanding of topic" = "understanding a topic based on what you're already familiar with". Explains why quantum mechanics doesn't feel "intuitive". There is nothing in our experience that makes the phenomena "familiar" to us.
But when my brain tries to learn math I haven't seen before, I usually can't extract intuition from the formula alone (compare this article with Wikipedia math pages). Short of playing around with the equations, I get a better understanding from articles like this when the topic is still new to me. There's less inertia to understand something like this, so I learn faster.
I don't think of math as a monster, but I appreciate simplification. The title is a bit click-baitey though.
But most people hate it that way. So I wonder if a better way to present statistics is to start with a random number generator and graphs. In fact, despite my love of equations, these days I haul out the random number generator anyway, to check whether my equations make sense.
Let me start off by saying that, as someone whose research is applied mathematics, I agree with you that math does not need to be "scary" if the pedagogy is tailored to the right level and style.
That being said...I think you're being a bit optimistic (or cavalier?), and maybe swinging too far to the other end of the spectrum. I don't believe programming has much overlap with mathematics at all. It's very close to applied logic, but I think even then it's specifically much more like an engineering discipline than a mathematical one. At the higher levels of computer science (like complexity theory) I see a much broader overlap with mathematics, but you use very little of that in typical programming.
I think the skills that make someone a good programmer and the skills that make someone a good mathematician are essentially orthogonal: I've met mathematicians who are excellent programmers (and vice versa), but I personally think that's more because they have learned to keep one discipline "out of the way" of the other one. I don't think most computer scientists or mathematicians actually make very good programmers (and vice versa) because there is little actual overlap in their skillsets.
Once you get past calculus (and maybe linear algebra), math becomes primarily about proofs, not computation. Proving abstractions in math and programming software do not have a significant overlap. I think most people can reasonably learn an undergraduate math curriculum (say, up to abstract algebra and something like elementary number theory), and I think most people can reasonably learn programming. But I don't see any realistic skill transference between these two, i.e. learning one won't make you learn the other any faster.
I'm not trying to claim one subject is obviously harder than the other, I'm just saying that it might give someone false confidence to have them dive into complex math just because they have experienced success in learning new programming frameworks or paradigms every year for their career. That is setting them up for failure - it takes significant mathematical maturity to read through a textbook and understand it without a teacher, just as it takes some domain-specific maturity to optimize learning a new programming language from a textbook or documentation.
People who want to learn math should be prepared to spend 5 - 10 minutes reading each page of a math textbook to really understand the material, followed by a few hours on each chapter's problems. This is absolutely doable for an autodidact, but I think my framing is a bit more realistic than yours. Even if they're innately capable of learning the math, they should be prepared for it to be essentially as difficult as learning programming for the very first time.
Then there's the naming ... P(X) and we're in probability and X is a random variable but if we are using just "p" then we're in physics and talking about momentum. If it's "P" then it's power, unless it's P(x) and then we're talking about geometry. And then you can italicize it, bold it, put a hat or a dot on it or under it, make the braces square, straight, or curly, and you get something else entirely; I mean completely different fields. Sometimes the same thing is used in multiple fields and then you need substitution syntax when using the two together. Great system!
So if you have a job where you are juggling 6 or so disciplines and you see a jumble of stylized letters in an equation along with a description, it's a fairly absurd system especially when the author assumes that all the readers know what they mean when they say i+p(j)/k or that when they use common everyday words which, in that particular branch of mathematics are actually very specific technical terms.
I often read things and think "what on earth is this person saying?" and then have to go back and decrypt this terribly designed mathematical language everyone uses that we are all supposed to say is a glorious and perfect interface. It's not, it's god awful and horrendous. The vast majority of humanity run screaming from it and can't interface with it to work through even simple concepts that they probably already know.
It's the modern version of ancient latin and it's become equally heretical to insist that we must create a better, more humane, more consistent, more discoverable, more flexible interface to describe the world that isn't just vestigial symbols from far-flung authors spanning 3000 years thrown together in a huge dumpster fire.
Mathematics is only different because it is useful in many other fields, so you get lots of non-mathematicians trying to make use of some result in isolation. When they don't understand the explanation, they blame it on the unfamiliar words and notation, but often that's just a symptom of not understanding the concepts. Anecdote: I have been taking a course taught in Chinese, which I don't speak very well, so all mathematical jargon is new to me. But just seeing how they were used let me recognize the words for familiar mathematical concepts.
Of course some mathematical writing is just bad, but usually mathematicians will then agree that it's bad. But if some formula comes with an explanation using words you don't understand, there is no way around looking up definitions until you understand the prerequisites. There is no language you could translate mathematics into to make it magically more understandable, unless the translation always prepends an introductory textbook.
I don't think it's heretical to long for better notation, I just don't think it's going to help much. But everyone is free to make up their own symbols, so feel free to go ahead. If it's really better, people will adopt it, just as they adopted superscripts for powers instead of one-off symbols, Leibniz notation instead of Newton's, and so on.
As a kid, I tried to keep away from math not because I hated it but because my math teacher did not know how to teach it. It was my conclusion then that they had not learned the subject thoroughly. As an adult, I decided to take matters into my own hands and started learning it by myself.
This happens a lot here, focus is lost very easily.