Bayesian updating of Probability Distributions
databozo.com
databozo.com
Someone with open mind has a priori with at least slight probability assigned to unlikely (for them!) hypothesis while on the other hand very religious people for example have 0 in their priori when it comes to possibility of their religion being made up so they are forced to ignore evidence to the contrary (because bayesian updating breaks for them due to division by zero and mind's way to signal this exception is denial). In general someone with a lot of weight on given hypothesis is "stubborn" or just very convinced and someone with uniform or close distribution just doesn't know anything about given problem.
Someone unable to build heavily weighted distributions is a conspiracy theorist, someone reluctant to - a sceptic and someone too much eager to a fanatic. Someone with very bad priors is un/badly educated (in given domain) or biased or maybe just stupid, someone with good priors is an expert. It's possible to combine expert with sceptic attitude or expert with fanatic or all too often stupid with fanatic (very bad and very heavily weighted priors with possible 0's on some options).
Once you start thinking this way you start expressing yourself differently, you start adding those probability qualifiers to your sentences: "I am very sure it's the way to go", "My intuition tells me this but I am not really sure", "I am very convinced and it's not worth discussing" (yes, it can be rational and good attitude) or "I would do X but I need more evidence to be reasonably sure".
It's all there in people's mind, language and interactions once you start thinking this way it's whole new world of perspective and understanding.
Why was the uniform distribution on [0, 1] chosen initially? Choosing a different distribution would give a different result. (And it doesn't make much sense to say, "Always choose the uniform distribution!" because the choice of variable affects the meaning of the distribution -- if instead we wonder about the value of p^2 and choose a uniform distribution for it on [0, 1], won't we get a completely different result?)
-It's a special case of the beta distribution, which is the conjugate prior for binomial problems. This means that the distribution of the probability of getting heads given the coin flips is in the same family as the prior itself (ie: beta priors with binomial likelihoods yield beta posteriors).
-The uniform (for this problem at least) is an "objective prior", which expresses that we don't have much information about whether the flip is biased. The example you give (modeling p^2 instead of p) is a great example of when the uniform would be a bad choice. The reason the uniform doesn't work in this case is because for binomial data (coin flips), a uniform prior is not invariant to reparametrization.
If choosing priors was so simple as always going with the uniform, there'd be little reason to go with Bayes! The choice of prior sometimes makes a radical difference in the posterior (especially with small samples), and there's many things to consider when you choose priors (computational convenience, uninformative versus informative priors, hierarchical modeling, etc).
http://en.wikipedia.org/wiki/Jeffreys_prior http://en.wikipedia.org/wiki/Beta_distribution http://en.wikipedia.org/wiki/Conjugate_prior
That may conflict with your sensibilities if you expect very strange coins, but if the priors aren't too crazily bad, if you have enough time for a few more coin flips, and the coin isn't particularly brittle, what's an order of magnitude between friends?
Edit: And for formal mathy purposes, ask someone else :)
Because physics - it's not possible to bias a rigid body so that it rotates with non-constant angular speed when flipped, as long as air resistance can be neglected, and that means that a fair flip gives 50/50 odds as to what side you catch it on. (Edit: clarity)
If other stuff is going on, like you're letting it bounce or something, then it depends on the particulars. It's rather easy to load a die, for instance.
From that perspective the uniform prior I use is reasonable (though non-committal and fairly uninformative), a triangular distribution with its peak at 50% (http://en.wikipedia.org/wiki/Triangular_distribution), the prior I have after updating tails and then heads, a normal distribution, or using the beta distribution with an equal number of both parameters.
Honestly you may decide there's a fair chance the coin is heads biased so you choose a prior that has most of the mass of the probability above 50%. As long as you have a reason, it's reasonable. You don't have to worry too much that you've chosen the optimal prior since more data makes equals of priors in the long run.
The only terrible thing you can do, is pick a prior that absolutely excludes certain hypotheses with 0% likelihood. No amount of data can overcome that.
The distribution for a coin is uniform over the event space [heads,tails]; not over the angle of rotation it lands on. You can have continuous or discrete event spaces.
If you have something shaped like a gömböc[1] you can have a "coin" that lands on heads ~100% of the time, except for the whole not-actually-being-a-coin part.
With just the few updates I've given you're probably right that it would affect things significantly. However, the more data you have the less the prior matters. This is known as swamping the priors.
In this case a uniform prior isn't incorrect, but you could definitely say its suboptimal and that I could make my examples much more accurate by choosing a prior that most represents my initial beliefs (like the one head one tail histogram for example.
In this context, at least, the prior distribution encodes everything - there's no meaning to the idea that you're more or less confident in the prior, because the prior already represents your uncertainty about the outcomes. If you were 50/50 on this prior versus another one, then your actual prior would be the average of the two.
There is a subfield of statistics that deals with imprecise probabilities, but that's a whole other can of worms and doesn't really relate to this problem. That said, it's fascinating stuff, and very useful in some contexts (if you're uncertain about your priors, it can be useful to do sensitivity analyses to figure out exactly how the end result depends on your prior).
Your explanation has the benefit of being a better explanation; my goal was just to explain where the inertia was hiding.
http://web4.cs.ucl.ac.uk/staff/D.Barber/pmwiki/pmwiki.php?n=...