Conditioning Is Grouping By
alexmolas.com
alexmolas.com
I personally prefer the "area of shapes in a plane" model. The sample space is a rectangle with area 1. Events are shapes inside the sample space, with probability equal to the % of the area of the sample space covered. The probability of intersection of two events is the area of the intersection of their corresponding shapes, and likewise for unions.
Conditioning on B is like treating B as the whole sample space. So what's the conditional probability of A given B? It's just the % of B that is covered by A. How much of B is covered by A? Of course, that's just the intersection of A and B. Then we divide by B to rescale that area relative to B, rather than the original sample space.
This "group by" model says something similar of course, but the "area of shapes" model I think covers a more general case while also being a literal mental picture that just about anyone can hold in their head.
I think a conventional way of describing this, which I think is in intro textbooks, is better -- we thinking of cutting slice through a multivariate distribution, and then normalizing. If you have some distribution P(X, Y) where X and Y are both real, then its density is like some landscape (i.e. each x,y is associated with some height). P(X | Y=y) is like a slice through that landscape at Y=y, which leaves a 1-D plot of the density along X. If we re-scale it so that the area under it is 1, then it's a well-formed PDF.
How intuitive!
/s
And this wikipedia page on iterated function notation https://en.wikipedia.org/wiki/Iterated_function
You'll also probably like Think Stats by Allen Downy which approaches teaching stats in a very similar framework (that you can code yourself) https://greenteapress.com/thinkstats/
It goes from a value x to a probability distribution of Y.
It is more general than a non-deterministic computation. You get a list of results with weights attached to it. (ignoring continuous distributions)
Grouping can be regarded as a non-deterministic computation.
It goes from a value x to a list of Y.
Computing averages or other aggregations from lists or distributions is another step.
https://dennybritz.com/posts/probability-monads-from-scratch...
https://github.com/mattearnshaw/lawvere/blob/master/pdfs/196...
Can you elaborate on that? I don't see how grouping can be a non-deterministic computation. When you compute the groups of a set you always get the same ones, or am I missing something?
The groups are deterministic. The non-determinism lies in which element of the group is then selected.
"Calculate an even number"
equals
"Take any element from the group of even numbers"
https://www.schoolofhaskell.com/school/starting-with-haskell...
Anecdote time: some weeks ago I had a bug for this reason. I was doing a groupby + select the first element, and since the first element is not always the same I was getting some weird results. I solved it by sorting the group elements before selecting the first element.