The problem with viewing these things as mathematical problems is that you buy into the simplifications you've made to fit the task to the model, and stop thinking about how the thing is actually working in your specific case. In short, it's a leaky abstraction.
Convergence is a good example. Lots of machine learning courses will go through the convergence proofs of say, the perceptron algorithm --- even though in all the hard problems, it's probably a really bad idea to let your algorithm converge! Your features are high dimensional and noisy, and so are your labels. Any linear separation you find is going to be terribly over-fit.
Taking it back to the blog post, when thinking about this k-Means clustering, it's actually much more important to think of it in terms of, "okay, all the algorithm gets to see is a term-document matrix". You should then expect that most other operations over term-document matrices will probably give you similarish results.
The difference between, say, k-Means clustering, and a LDA model, are still important for someone working in the area to understand. They should also understand how a hierarchical dirichilet process solves problems you might have for an LDA model.
Ultimately, you need to be able to develop a good tool box, and that means reasoning about the performance of various methods on your problem domain. (It doesn't mean "blindly twiddle nobs on a collection of crappy implementations, ala weka...")