I don’t necessarily think that an article on math should cover everything that is known about a subject. It’s okay to write for an audience that doesn’t know anything about autocorrelation and multi-chain methods. The need to elaborate so thoroughly on every possible prerequisite and/or application, and cover every corner case you might have is one of the reasons so many people dislike reading math on Wikipedia ... it’s unapproachable unless you already know everything about it. It’s becoming a reference only for experts not very usable for learning by a student.
Anyway, what’s the true danger of a simple or incomplete understanding? It’s not that likely to lead to people putting the wrong thing into nuclear reactors; isn’t it more likely to lead to someone getting the wrong answer in a weekend project software and then spending Sunday learning a little more about MCMC methods?
In practice the danger is someone not trained in statistics will copy/paste the “simple” approach and generate poor chains of samples, and base a seriously incorrect MCMC calculation off their misunderstood application.
If it’s clear this is just for teaching, then sure the risk is less. But it’s not usually clear.
I think explanations don't require formal proofs to be useful.
Otherwise you are cargo cult copying some code and expecting it to work, without understanding the algorithm the code is executing.
Further you need to understand the convergence aspects too, because when you realize a real sample in a chain from your software application, whether or not you can safely use that chain of samples for the estimation or simulation you intended really depends on autocorrelation and convergence criteria. You cannot just believe the code was OK so the samples can be used... but “intuition” tutorials like this give a false sense of security that you can.