Improbable Inspiration: Bayesian Networks (1996)
cs.ubc.ca
cs.ubc.ca
The "mathematical equations by Danish researchers", for those interested, are most likely this paper:
Lauritzen, S.L. and Spiegelhalter, D.J. (1988), Local Computations with Probabilities on Graphical Structures and Their Application to Expert Systems. Journal of the Royal Statistical Society: Series B (Methodological), 50: 157-194. https://doi.org/10.1111/j.2517-6161.1988.tb01721.x
Direct PDF Link: https://www.eecis.udel.edu/~shatkay/Course/papers/Lauritzen1...
For example, Pyro implements tons of facilities to have Bayesian models augmented with neural networks.
It makes a lot of sense from a modeling perspective to model the big picture using a Bayesian model (generally a graphical model) and then use neural networks for some components. You capture the overall causal structure, but you are also outputting really precise predictions. For example, a deep markov model.
There are tons of unexplored ideas combining both, and in general I think this is the future of deep learning and one component towards AGI.
The other main issue is that in graphs with multiple paths to a single node, you can’t quite do this forward backward pass and get the _exact_ answer. You can approximate it by just passing messages in these cycles that arise, but you’re not actually guaranteed to converge to the right marginal probability anymore. This is called Loopy Belief Propagation. Additionally, there’s no real true ordering of nodes for a passing order. There’s not a natural sequence like there is from leaves to root and back when you can go around and around in circles.
It surprisingly works reasonably well in a lot of cases anyway though. BP and the various approximate versions are super interesting. The original algorithms really only work on low dimensional discrete spaces or things with analytic solutions to integrating from conditional/joint to marginal distributions. However, there’s been some really cool stuff coming out in the last 5ish years about using particle based approximations to work on more complicated/continuous spaces.
I think maybe reinforcement learning where human feedback becomes part of the loop is about as close as I could think of. But that is different than factoring in human input to probability calculations.
Any time you see graphical models, they’re usually BNs. Undirected graphical models are very closely related too (all directed models can be represented as undirected models, but not all undirected models can represent directed models), but they’re usually not referred to as BNs.
They’re used all over the place. One school of causal inference is heavily steeped in BNs/DAGs. This shouldn’t be surprising because the creator of BNs, Judea Pearl, is heavily involved in causal inference now.
> The framework estimates the climate change risks in economic terms by modeling the main activities that a mining company performs, in a probabilistic model, using Bayes’ theorem. The model permits incorporating inherent uncertainty via fuzzy logic and is implemented in two versatile ways: as a discrete Bayesian network or as a conditional linear Gaussian network. This innovative quantitative methodology produces probabilistic outcomes in monetary values estimated either as percentage of annual loss revenue or net loss/gains value.