Judea Pearl, big brain behind AI, wins Turing Award (Nobel Prize in Computing)
networkworld.com
networkworld.com
For HNers just out of college, it's important to note that Bayesian analysis has not always been as popular as it is now. In fact, even as recently as the 1990s, it was regarded with suspicion by many statisticians, who strongly disliked the idea that prior and posterior distributions are meant to represent subjective states of belief. I was fortunate enough to have a very progressive statistics professor in undergrad in the 1990s, who was interested in Bayesian analysis. It's my understanding that most upper-level statistics and probability coursework avoided doing much Bayesian analysis until around the turn of the millennium. (if you went to university in the 1990s or earlier, was this your experience?).
For those interested in learning more about this topic, an excellent book on the history of the Bayes theorem controversies was recently published by Yale Unviersity Press: _The Theory That Would Not Die_ by Sharon Bertsch McGrayne
To be fair to Judea Pearl, I do see that he wrote an essay entitled, "Bayesianism and causality, or, why I am only half-bayesian". Nonetheless, much of his work appears to involve what we now know as Bayesian analysis. So like many (all?) scientists who make major breakthroughs, Judea Pearl was going against the accepted understanding of probability by pushing Bayesian analysis in the 70s, 80s, and 90s. It's good to see his daring rewarded.
I don't see Pearl as primarily interested in the Bayesian v. frequentist debate himself, though, but rather in how to efficiently do probabilistic reasoning in non-trivial problems in general, with a heavy tilt towards questions of representing causality. Methodologically his work over the years has used all sorts of things from various camps; for example, he was also an authority in the early 1980s on heuristic search.
This is HackerNews! People here are expected to know what the Turing Award is; if they don't, they can find out for themselves.
Maybe there is no harm in saving people from having to do a search, but there is harm in being misleading.
[1] http://www.networkworld.com/news/2011/060611-nobel-prize-com...
Search belief propagation, junction tree algo, markov blanket, belief propagation statistical physics. The last search showing his ideas finding use beyond where he introduced them. Whenever ideas keep showing uses in different places especially in something as well evinced as the thermodynamic/information link then you know you have done something profound.
He is also credited with first describing the ever-popular belief propagation algorithm for approximate reasoning in Bayesian networks and graphical models in general.
Edit -- I found this to be a nice, short exposition: http://ftp.cs.ucla.edu/pub/stat_ser/r368.pdf
The idea of Markov chains predates Pearl. His work was a demonstration of the accuracy and power of Bayesian inference networks. He revitalized that idea at a time where programmers were just beginning to have the data and processing power to apply machine learning.
The rest is history.
Post hoc ergo propter hoc? The success of naive Bayes classifiers for spam filtering must have improved popular perceptions of Bayesian techniques, but let us not go overboard.
Bayesian Networks tend to be used not as classifiers but a tool to explore joint probability distributions.
Interestingly related to your topic, HMMs and Naive Bayes are related in that HMMs are kinda like the sequence sensitive version. HMMs and Naive Bayes are generative models. They both model/estimate a joint probability on the data with very strong conditional independence assumptions. Where as Logistic regression and Conditional Random Fields estimate the conditional probability of the output/labels directly.
HMMs : Naive Bayes as linear chain CRFs : Logistic Regression. CRFs are state of the art at sequence and time series prediction. I have not yet gotten my head round them though. The relationship between logistic regression and naive bayes is not commonly known (although the comparison of log reg to a simple Neural network is common). Knowing when Logistic regression outperforms Naive Bayes is useful (simple rule of thumb: logistic regression less sensitive to independence assumption, more data use log reg, less data use naive bayes). I've implemented a multi class sparse regularized logistic: SMLR. Its up there with linear SVMs but simpler but also gives a probability.