First, Bayesian formalism shouldn't be hard. Just remember that
P(A | B) = P(A & B) / P(B)
and it's easy to derive (left as exercise) Bayes's Law. That's quite easy. The hard part, here, is applying it.Let's say that I have a coin and I run a game where I always bet heads. If I flip tails, you get $1; on heads, I get $1. Let's also say that 1.00% of the people in the world are depraved enough to use unfair coins that always come up heads. You don't know me, so that's a good assumption for how likely I am to be a scumbag. So there's a prior probability of 0.01 that I'm a scumbag running an unfair game, and 0.99 that I'm using a fair coin. Now, we play, and I flip a head. I win. There's evidence (or signal) that is suggestive that it's more likely that I'm a scumbag, but far from conclusive (fair games will turn up heads as well). How much has the probability changed? Well, let's look at the possibilities:
I'm a scumbag, and flip heads: 1/100 * 1 = 1/100
-- There's a 1% prior prob., and if I'm a scumbag I'll always flip heads.
I'm a scumbag, and flip tails: 1/100 * 0 = 0
-- If I'm a scumbag, I'll never flip tails.
I'm using a fair coin, and flip heads: 99/100 * 1/2 = 99/200.
I'm using a fair coin, and flip tails: 99/100 * 1/2 = 99/200.
It might help to draw a rectangle for the four possibilities, like so: Scumbag Fair
(1/100) (99/100)
+---------+---------+
| Scumbag | Fair |
| Heads | Heads |
| | |
| 1/100 | 99/200 |
+---------+---------+
| Scumbag | Fair |
| Tails | Tails |
| | |
| 0 | 99/200 |
+---------+---------+
Once we flip a head, we can throw out the "flip tails" sections, because we didn't. We flipped heads. That leaves us with a subspace that contains 99/200 + 1/100 = 101/200 of the total space-- we know we're in that space, so we can consider only it-- while the probability mass of the "I'm a scumbag and flip heads" remains 1/100. So the probability that I'm a scumbag posterior to flipping heads is: (1/100)/(101/200) = 2/101 ~ 2%
Probability wise, I'm almost twice as likely to flip a coin after a head comes up. That relationship holds very well for small priors, but what if there's a 55% chance that I'm a scumbag? Flipping heads doesn't put that probability to 110%-- first of all, that makes no mathematical sense, and second of all, there's still a chance that I'm fair but just flipped another head. (The actual posterior probability is about 71%.)For probabilities over 0.5, "twice as likely" doesn't make sense, if you think in terms of probability. What about odds, however? Then, you can derive that "twice as likely" as 55% is 71%.
You can say "twice as likely" if you think in odds, not probability. Odds is probability transformed through p/(1-p), or the ratio between the probability of the event happening and it not happening; it has domain [0, +inf] which means that "twice as likely" doesn't cause that problem. If you think in odds terms, it turns out that odds(I'm a scumbag) exactly doubles each time I flip a head. After the first, that odds number goes from 1/99 to 2/99. If I flip 10 heads, then it's 1024/99; transformed back into a probability it's 1024/1123, or a 91% chance that I'm a scumbag.
In fact, what I think makes Bayesian inference hard is that it involves that subtle context switch between probability and odds.
If B is observed and there is a known prior _odds_ of A, the posterior _odds_ of A
are k times higher, where k is the ratio of the _probabilities_ of
(A & B) vs. (~A & B).
What makes Bayesian inference neat is that you can run it as an online process (you can look at events one at a time). If observed events (signals) are independent, you don't need to process the corpus as a whole, and order doesn't matter. That becomes nice when you factor in concurrency and distributed computation. From these principles, you can also derive logistic regression (multiplication in the log-odds space is a vertical shift along an S-shaped logistic curve) and you start to see why the logistic curve comes up so often in machine learning.Ok, so what about Bayesian statistics? Essentially, it comes from the idea that:
(1) There is some prior probability distribution [0] over the possible states S
of affairs, which you cannot observe directly but whose relative
probabilities you can estimate based on observed events E.
(2) For each event E observed, compute unnormalized posteriors according to
posterior_unnormalized(S) = prior(S) * likelihood(E | S).
(3) To compute normalized posteriors, divide each of the unnormalized ones
by the sum (or integral) of those likelihoods over all states S.
You now have a sum of 1, which makes it a legitimate probability distribution.
[0] Regarding (1), picking priors is more of an art than a science, which is why some people distrust Bayesian methods. The good news is that if you have a lot of evidence and your priors are reasonable (few assumptions / Ockham's Razor) you will converge to something close to the right answer with enough evidence, regardless of priors.Now, many actual Bayesian inference methods ignore (3). It's computationally expensive (when there aren't two possible states S, but millions to infinitely many) to normalize and often we don't need to do it. For example, if you're doing classification between two classes Q and R-- with "events" being features of what you're trying to classify-- then you generally only care about which is more likely, not whether there's specifically a 23.75% chance of Q. So the normalization step is generally either skipped (if relative probabilities are all that matters) or procrastinated to the end of the analysis. This "feels wrong" at first but it actually works, and is often better (numerical stability) for the accuracy of the conclusions.