Birthday Paradox Revisited
datagenetics.com
datagenetics.com
How many people do you have to put in a room
before there's a 50% chance of two having the
same birthday.
Assuming a uniform distribution of birthdays. The answer turns out to be about 23. In general, for M possible birthdays (other than 365, say), you need about sqrt(M) people before you get a 50% chance.The article is a standard introduction to the Birthday Problem with some real world data thrown in. The 'paradox' comes from the surprise that the number is so low (23) compared with the number of days in the year (365). As the article points out, one short description of why the number is so low is that you're comparing each new person with every person already in the room instead of drawing two numbers at random and seeing if they're the same.
For the curious, I have a minimal post on how to derive the Birthday Paradox and other canonical probability problems [2].
[1] https://en.wikipedia.org/wiki/Birthday_problem
[2] https://mechaelephant.com/dev/Assorted-Small-Probability-Pro...
Now, if it were impossible to ever get multiple matching pairs, then "the chance there's at least one match" would be equal to "the expected number of matches". Specifically: "chance of at least 1 match" = "expected number of matches" - "chance of at least 2 matches" - "chance of at least 3 matches" - ... . Since it's possible but unlikely to get multiple matches, the approximation should be reasonably close.
Even thinking about this more, and reading (and understanding) the solution, if I think about this, my guess would still be "well, with 182 people, half the pigeon holes would be taken, so there would be 50% chance the next person getting in the room would pick a taken hole".
Another similar problem is the Monty Hall problem. Simple, easy to understand when explained, but still, despite understanding the solution, doesn't feel right!
Its true that if you have 182 people in the room all with unique birthdays and you add one more random person to the room there is a 50% chance of them sharing a birthday just like you described. But you're assuming you already managed to gather 182 people without any birthday collisions. That's a different problem then the one originally posed. In your case you collected 182 people with unique birthdays and checked the probability of a collision when adding one more. But the real question asks what is if you grabbed those 182 people at random what is the chance that any two of them already share a birthday?
However I did just read about Bertrand's Box Paradox[1], and it's very much the same sort of thinking as the Monty Hall problem, but more intuitively understandable for me at least.
[1] - https://en.wikipedia.org/wiki/Bertrand%27s_box_paradox
Does this make any sense? The author says that with a realistic non-uniform distribution of birthdays, they found that collisions were less likely when there were fewer than 5 people involved. I can't think of any reason this would be the case. If certain birthdays are more common than others, I'd think that the chances of a collision must go up regardless of the number of people involved.
Is there any plausible mathematical explanation for this effect? Or was the experiment just underpowered? Or worse, might the simulation code be buggy? Presumably they would have run the simulation more than once after getting such such a counterintuitive answer, and it seems really unlikely that this effect would be consistent unless something was broken.
At least for the case of two people, if we say that a_n is the probability of a person's birthday being chosen (so that ∑(a_n) = 1), then the chance of a match is ∑(a_n^2). To construct a rigorous argument, one could bring in the power mean theorem[1], which states that, if x > y and the sequence {a_n} is all nonnegative, then (avg(a_n^x))^1/x ≥ (avg(a_n^y))^1/y, with equality if and only if all elements of {a_n} are the same. This can be used to show that, if ∑(a_n) is fixed, then ∑(a_n^2) (i.e. the probability of a match) is minimized only when {a_n} are all equal.
It's less obvious how to make a rigorous argument for the cases of 3 and more people, but https://en.wikipedia.org/wiki/Muirhead%27s_inequality might help (if you trust the proof of it).
[1] https://en.wikipedia.org/wiki/Generalized_mean#Generalized_m...
The author suggests such a reason: the loss of precision. As I understand it hints to a limited precision of float/double number representation. When you calculate a sum of millions of fractions some of fractions could turn into zero, because you have divided a too small number on a too large one.
But presuming the author ran the simulation more than once (wouldn't you do this if you had such a surprising result?) I don't see how there could be any consistent effect unless there was a bug in the logic of the program. Unless maybe they ran it multiple times with the same random seed and got exactly the same results?
The interesting thing is that induction increases the risk of c-section, which is a negative for the family, but a positive for the healthcare system (for obvious reasons).
Also note that the US has one of the worst birthing mortality rates amongst western countries...
USA#1
My father and step-sister also share their birthdays.
I guess probabilities are just really unintuitive to people.
"You know, the most amazing thing happened to me tonight... I saw a car with the license plate ARW 357. Can you imagine? Of all the millions of license plates in the state, what was the chance that I would see that particular one tonight? Amazing!"
I know this because our son's due date was the 1st of May and the staff was explicit about this.
My older sibling was a C-section, so I was going to be a C-section anyways. It just became a planned c-section due to the holiday.
Not as many mothers (or doctors) in the U.S. seem to want to induce labor on U.S. Independence Day.
A seven-day smoothing realy reveals this.