An Introduction to Model-Based Machine Learning
blog.dominodatalab.com
blog.dominodatalab.com
I transitioned from using Bayesian models in academia to using machine learning models in industry. One of the core differences in the two paradigms is the "feel" when constructing models. For a Bayesian model, you feel like you're constructing the model from first principles. You set your conditional probabilities and priors and see if it fits the data. I'm sure probabilistic programming languages facilitated that feeling. For machine learning models, it feels like you're starting from the loss function and working back to get the best configuration.
Much of the underlying machinery behind Bayesian vs. machine learning models is the same. Hidden Markov Models are Hidden Markov Models whether they have a prior or not. But this difference in feel influences how you build models and hence, the results.
Now that optimization algos for Bayesian models are catching up, Bayesian ML might become a thing.
Cool stuff.
In the literature, Bayesian models are usually considered a type of machine learning model.
However, compare the topics covered in Duda and Hart's text (1st edition, 1973; 2nd edition 1995) to the more recent text by Hastie et al. (current edition ~2013). There isn't a huge difference in subject matter. The latter is slightly more advanced and has more of a statistics perspective, but the foundations are there: Bayes' decision theory, linear methods for classification / regression, naive bayes, neural networks, decision trees, ensembles, clustering, etc...
There is a range of topics that are foundational to ML, and thus, relatively stable. These topics are built upon the even more solid foundations of probability theory and statistics. The biggest advances in ML in the last decade (I would argue) were not due to advances in theory.
Is that 'The elements of statistical learning'?
If we knew the right answers we wouldn't waste our time implementing the wrong ones. It's not a bad thing: it means there is room for you to discover something genuinely new and understand something, however small, that nobody else ever completely understood.
I think that's great. But YMMV.
Finding a book that hit the sweet spot for regressions wasn't easy but was doable. I was hoping there would be something similar with Bayesian Nets/Generative Models.
Could you provide an example?
This is a book that emphasizes practical applications without getting bent on the math details too much. If on the other hand you are a math whiz, Elements of Statistical Learning is THE book but it expects you to be very proficient in math.
Both books are seriously underrated, which is kind of funny to say because you will find only praises about them, but they deserve even more.
If the math in ESL gives you trouble, you might prefer http://www-bcf.usc.edu/~gareth/ISL/ISLR%20First%20Printing.p... (and I'm not just saying this because the first author was one of my advisers, although I do think that he and Daniela are particularly gifted teachers).
If the math in ESL is too trivial for you, there's always https://web.stanford.edu/~hastie/StatLearnSparsity_files/SLS... , which covers some graphical modeling strategies in later chapters and even kicks the tires of the autoencoder (imho perhaps the greatest recent advance in neural networks for practitioners) along the way.
Koller's course and Ng's course are also good.
Ultimately I feel like you have to get the math right or you'll never acquire the intuition that helps you design your own approaches. But you also have to put in the work.
That reminds me, tibshirani's Stanford course (accompanies ISL and ESL) is terrific. Better than those other two, actually. I wish Harrell would offer one.
>Ultimately I feel like you have to get the math right or you'll never acquire the intuition that helps you design your own approaches.
But what do you mean by that? Do I really need it if my applications are not as demanding as Netflix? I feel like many people consider anything less than phd-level understanding lol worthy, which is simply not true. Majority of analysts out there are doing just fine with canned procedures. Are there something like canned procedures for Bayesian Nets?
Re: "do I really need it?": hell if I know, I'm not you. But my assertion was specifically that if you want to design your own methods (i.e. do research) you need to understand what they are doing. This doesn't seem like a controversial position; an expert is simply a master of the fundamentals.
Linear algebra and calculus (to a lesser degree) are foundational for a great many things. Got missing data? K-NN or nuclear norm matrix completion (or marginalizing over the rest) can help. Systems of differential equations? Use a matrix exponential.
You are free to do whatever you like. A bus driver doesn't need to know how to rebuild an engine. But if you want to race cars you'll get a lot further if you do know how.
So after all this back and forth, I still don't know if there is a book similar to RMS in scope, for Bayesian Nets.
> 1. This approach provides a systematic process of developing bespoke models tailored to our specific problem.
> 2. It provides transparency to our model as we explicitly defined our model assumptions by leveraging prior knowledge about traffic congestion.
> 3. The approach allows handling of uncertainty in a principled manner using probability theory.
> 4. It does not suffer from overfitting as the model parameters are learned using Bayesian inference and not optimization.
> 5. Finally, MBML separates the model development from inference which allows us to build several models and use the same inference algorithm to learn the model parameters. This in turn helps to quickly compare several alternative models and select the best model that is explained by the observed data.