Is there a good "bayesian probability for dummies" type resource out there? Any suggestions?
Is there a good "bayesian probability for dummies" type resource out there? Any suggestions?
For example: I have two bags of sweets. One has 10 sweets, of which 8 are chocolate and 2 are marshmallows. The other has 20 sweets, of which 10 are chocolate and 20 are marshmallows.
I offer to give you a sweet, which I will select at random. Assuming I empty both bags on to the table and choose one at random (uniformly), what is the chance it comes from the first bag? Assume that I do this where you can't see what I choose.
(Answer: 10/30, or 1 in 3 - as these are the ratios of total sweet numbers).
Now, assume that after I've chosen (you still can't see), I tell you that the sweet is a chocolate (This is the evidence). What is the chance now that it came from the first bag?
(Answer: 8/18 - the ratio of first-bag-chocolates out of any-bag-chocolates)
We have now 'updated' our probability estimates based on the evidence ('It's a chocolate'). Important to note here is that the first bag has a higher ratio of chocolates than the second bag, BUT it's still more likely to have come from the second bag due to the second bag having more sweets in total - what we call the Prior probability.
This relationship can be expressed mathematically, which is what Bayes did.
P(B|A) = P(B & A)/P(A) = P(B)*P(A|B)/(P(B)*P(A|B) + P(!B)*P(A|!B))
(Where P(X) means The probability of X occurring, P(X&Y) means the probability of both X occurring and Y occuring, and P(X|Y) means the probability of X occuring, given that Y has already occurred (or that we assume it occurs))Also, a good idea is not "trying to turn your mind inside out", which is what seems to be needed from a lot of explanations I have read.
It is easy, it is simple, it is normal.
I am sorry I do not have a reference handy (because of the above) but those are good suggestions (otherwise I would not be writing them) that have helped me overcome similar "difficulties".
Edit: Cannot help saying it. The 'a priori' and 'a posteriori' and all those terms are just fancy words (which mean what they signal) which scare a lot of people away from something, I repeat, natural.
An Intuitive Explanation of Eliezer Yudkowsky’s Intuitive Explanation of Bayes’ Theorem
Seeing the world through the lens of Bayes’ Theorem is like seeing The Matrix. Nothing is the same after you have seen Bayes.First, Bayesian formalism shouldn't be hard. Just remember that
P(A | B) = P(A & B) / P(B)
and it's easy to derive (left as exercise) Bayes's Law. That's quite easy. The hard part, here, is applying it.Let's say that I have a coin and I run a game where I always bet heads. If I flip tails, you get $1; on heads, I get $1. Let's also say that 1.00% of the people in the world are depraved enough to use unfair coins that always come up heads. You don't know me, so that's a good assumption for how likely I am to be a scumbag. So there's a prior probability of 0.01 that I'm a scumbag running an unfair game, and 0.99 that I'm using a fair coin. Now, we play, and I flip a head. I win. There's evidence (or signal) that is suggestive that it's more likely that I'm a scumbag, but far from conclusive (fair games will turn up heads as well). How much has the probability changed? Well, let's look at the possibilities:
I'm a scumbag, and flip heads: 1/100 * 1 = 1/100
-- There's a 1% prior prob., and if I'm a scumbag I'll always flip heads.
I'm a scumbag, and flip tails: 1/100 * 0 = 0
-- If I'm a scumbag, I'll never flip tails.
I'm using a fair coin, and flip heads: 99/100 * 1/2 = 99/200.
I'm using a fair coin, and flip tails: 99/100 * 1/2 = 99/200.
It might help to draw a rectangle for the four possibilities, like so: Scumbag Fair
(1/100) (99/100)
+---------+---------+
| Scumbag | Fair |
| Heads | Heads |
| | |
| 1/100 | 99/200 |
+---------+---------+
| Scumbag | Fair |
| Tails | Tails |
| | |
| 0 | 99/200 |
+---------+---------+
Once we flip a head, we can throw out the "flip tails" sections, because we didn't. We flipped heads. That leaves us with a subspace that contains 99/200 + 1/100 = 101/200 of the total space-- we know we're in that space, so we can consider only it-- while the probability mass of the "I'm a scumbag and flip heads" remains 1/100. So the probability that I'm a scumbag posterior to flipping heads is: (1/100)/(101/200) = 2/101 ~ 2%
Probability wise, I'm almost twice as likely to flip a coin after a head comes up. That relationship holds very well for small priors, but what if there's a 55% chance that I'm a scumbag? Flipping heads doesn't put that probability to 110%-- first of all, that makes no mathematical sense, and second of all, there's still a chance that I'm fair but just flipped another head. (The actual posterior probability is about 71%.)For probabilities over 0.5, "twice as likely" doesn't make sense, if you think in terms of probability. What about odds, however? Then, you can derive that "twice as likely" as 55% is 71%.
You can say "twice as likely" if you think in odds, not probability. Odds is probability transformed through p/(1-p), or the ratio between the probability of the event happening and it not happening; it has domain [0, +inf] which means that "twice as likely" doesn't cause that problem. If you think in odds terms, it turns out that odds(I'm a scumbag) exactly doubles each time I flip a head. After the first, that odds number goes from 1/99 to 2/99. If I flip 10 heads, then it's 1024/99; transformed back into a probability it's 1024/1123, or a 91% chance that I'm a scumbag.
In fact, what I think makes Bayesian inference hard is that it involves that subtle context switch between probability and odds.
If B is observed and there is a known prior _odds_ of A, the posterior _odds_ of A
are k times higher, where k is the ratio of the _probabilities_ of
(A & B) vs. (~A & B).
What makes Bayesian inference neat is that you can run it as an online process (you can look at events one at a time). If observed events (signals) are independent, you don't need to process the corpus as a whole, and order doesn't matter. That becomes nice when you factor in concurrency and distributed computation. From these principles, you can also derive logistic regression (multiplication in the log-odds space is a vertical shift along an S-shaped logistic curve) and you start to see why the logistic curve comes up so often in machine learning.Ok, so what about Bayesian statistics? Essentially, it comes from the idea that:
(1) There is some prior probability distribution [0] over the possible states S
of affairs, which you cannot observe directly but whose relative
probabilities you can estimate based on observed events E.
(2) For each event E observed, compute unnormalized posteriors according to
posterior_unnormalized(S) = prior(S) * likelihood(E | S).
(3) To compute normalized posteriors, divide each of the unnormalized ones
by the sum (or integral) of those likelihoods over all states S.
You now have a sum of 1, which makes it a legitimate probability distribution.
[0] Regarding (1), picking priors is more of an art than a science, which is why some people distrust Bayesian methods. The good news is that if you have a lot of evidence and your priors are reasonable (few assumptions / Ockham's Razor) you will converge to something close to the right answer with enough evidence, regardless of priors.Now, many actual Bayesian inference methods ignore (3). It's computationally expensive (when there aren't two possible states S, but millions to infinitely many) to normalize and often we don't need to do it. For example, if you're doing classification between two classes Q and R-- with "events" being features of what you're trying to classify-- then you generally only care about which is more likely, not whether there's specifically a 23.75% chance of Q. So the normalization step is generally either skipped (if relative probabilities are all that matters) or procrastinated to the end of the analysis. This "feels wrong" at first but it actually works, and is often better (numerical stability) for the accuracy of the conclusions.
I do wish every explanation of such things explained the syntax better. I still have to keep checking what P(A | B) and P(A & B) mean. You say the hard part is applying the equation, but the first "hard part" is learning what all the symbols mean!
http://jeremykun.com/2013/01/04/probability-theory-a-primer/
http://jeremykun.com/2013/03/28/conditional-partitioned-prob...
His next post in the series is supposed to be on Bayes' theorem.
http://blog.moertel.com/posts/2010-12-07-on-the-evidence-of-...