How Common Is Your Birthday? (2012)
thedailyviz.com
thedailyviz.com
What's confusing the visual are planned birth days.
Personally, there was a pretty significant blizzard ~40 weeks before I was born. Snowstorm babies unite.
Avoiding holidays was the obvious trend, but there were two others I saw that made me chuckle:
1. People avoiding February 29, so they don’t give their kid that 1-in-4 birthday pain.
2. The spike on February 14! People want their kid born on Valentines (cute, but not really romantic...)
I wonder if this error is common in other people. And if so, if there are social-network dynamics around birthdays that cause the bias.
https://www.bustle.com/articles/93018-can-your-birthday-pred...
If you mouse over the individual squares of the image that follows the words:
How Popular Is Your Birthday?
Two decades of American birthdays,
averaged by month and day.
you will get popups showing more details for each particular date.A sample popup text says:
this date,
10/2, had
11,572
births on
average. It
ranks 77th.
The
conception
date* is
around 1/9.At first I would simply correct and "educate" the end user how to interpret the charts. Recently though I have begun to "register" such types of questions and misunderstandings. I now actively determine WHY the end user made such a seemingly simple error in usage. And more than likely, I've concluded, it's the fault in the presentation of the chart than it is of the user's understanding of it. It is not easy to create data visualisations that do not mislead!
In this case, we as users are introduced to a gradual shading of a pink/maroon palette of calendar dates. The shading goes from dark (almost black) to light pink (almost white). So both white and black squares are tacitly presented as features. The mind extrapolates this. Using white for an invalid date is bound to create confusion to a user not paying attention. And charts are always sold to end users as "you relax, we've done the thinking, just consume this easy picture". As users, our guard is down by invitation!
I'd make all the invalid dates a colour from a completely different palette or strikingly different colour (eg. gray, yellow or green) to stand out clearly, indicating immediately that you are looking at invalid areas of the chart.
> 9/11 has a small but noticeable dip compared to the surrounding dates. I wonder if that's people avoiding it or if it's just coincidence
I also see something else: there is a light streak going down the 13th of every month, as if people were having fewer babies that day.
If I want to figure out whether this is an actual pattern or just a coincidence, how would I go about that? I mean, as I understand the general process it's a two-parter:
1. Come up with a model of the situation: say, uniform probability of giving birth any given day.
2. Calculate the probability of getting your actual results based on the model. If these are lower than some threshold you pick, say, 5% or 1%, then you discard your model because you have found something that breaks your assumptions.
But there's lots of nuance here. Is does my model even make sense? Isn't it more useful to define the number of child births on any given day to be normally distributed with mu = average over the year? Are the two equivalent because something something central limit theorem?
Is there a simple and convenient way to do the calculations for the most common models? And say I'm interested only in the streak of the 13th... do I need to consider all other days as well before I discard my model? My intuition says that the streak of the 13th is easier to prove/disprove than the dip of 9/11, because it's 12 data points (is there?) compared to just one -- but I also feel like I'm way off base here...
I'm at a bit of a loss here. I haven't studied a lot of statistics because I've been busy with other things but this is the kind of thing I'd love love love to learn.
I do remember a bit of the Think Bayes book, and I could plug each month into a comouterprogram updating the birth probabilities of each day based on that... but how do I know if I have enough data to get a day-by-day resolving power? Maybe I should look at weekly averages instead?
I mean, I understand this is a whole field of research and not something you learn overnight, but is there not some subset of "countertop hypothesis testing" that you can apply casually over the kitchen table to get you at least somewhere?
TL;DR: I think I see patterns in this data and I want to statistically reject my hypotheses. How do I approach that?
Edit: I guess the reason my question differs slightly from high school stats is because I don't get to design an experiment that will give easily analysable data. And I'm afraid of creatively interpreting the existing data to make it easier to deal with, because I think that will lead me to incorrect conclusions.
> The rate of caesarean section births in the U.S. was 32.7 percent in 2013
With so many c-sections being performed there is more selection of birthdates.
Anecdotally, I went in with my wife for an ultrasound on a Friday the 13th and it was definitely less crowded that day. Rest assured, no Rosemary's Baby detected, nor any heart defects, etc. I just benefited from not having to take a seat in the waiting room :).