Disproportionately Common Names by Profession
verdantlabs.com
verdantlabs.com
Shouldn't it say: "People with these names happen to be in those professions more often than others"?
Anyway, there are a couple of fun ones in there, but I'll let you figure those out yourself. Unfortunately neither my name nor my profession are covered — I'm not quite sure what to make of that. :/
---
They use the same language in their blog post: "Arnolds therefore appear to have a much higher tendency to be accountants than Shanes." That's just wrong, no? By wrong I just mean intentionally misleading. I'm sure that has nothing to do with the fact that they sell an app that helps you find names for you babies though.
Take above quote, which compares "1.9% of Arnolds are accountants" to the "0.55% of Shanes [are accountants]". They're implying that the probability of being an accountant (J), given that ones name (N) is Arnold, is above the expected probability of being an accountant in general. So they're looking for a high P[J|N]/P[J].
Now compare with what we were expecting to see. We assumed the chart showed, for a given job, names which had a higher incidence than normal. i.e., we're looking for a high P[N|J]/P[N].
Guess what. P[J|N]/P[J]=P[N|J]/P[N] by Bayes' Theorem [1]. These are EXACTLY the same metric! So their technique, and the chart, is correct. (And my original post, below, was wrong.)
(Not saying anything about causation here, and I don't think they were either.)
[1] http://en.wikipedia.org/wiki/Bayes%27_theorem
-----------
Yes, that is completely backward. 99% of Arnolds could go into farming, yet still the 1% who go into accounting dominate that field, and hence show up on this chart.
EDIT: but you missed the first half of that quote: "In our sample of two and a half million people, a whopping 1.9% of Arnolds are accountants. Contrast that with just 0.55% of Shanes." So I think the quote is correct. Makes me wonder if their chart is backward. (i.e., they put "Arnold" under "Accountant" because Arnolds are likely to be accountants, not because accountants are likely to be Arnolds, as the grouping implies).
Before we get to my confusion, one interesting thing I found along the way is the following situation, call J1 "job 1" and N1 "name 1", and use the P(N|J)/P(N) (or its equivalent) metric:
J1 J2
------
N1 N1
N1 N1
N1 N3
N2
If we limit each job to it's top name, N1 doesn't get attached to either despite being the most common name in each. J1 gets N2 while J2 gets N3.If this is the method then don't use this chart to guess the names of people in a profession, use it to guess the professions of people whose names you know. "Guy" may be listed for investment bankers, but an investment banker is still more likely to be named Dave, but if you meet a Dave he's likely to be a mechanic.
For the same situation say we use the top value for P(N|J), then J1 gets N1 and so does J2. P(J|N) goes back to J1 getting N2 and J2 getting N3 and N1 being left out in the cold.
But here's where it's unclear, I think this:
> In our sample of two and a half million people, a whopping 1.9% of Arnolds are accountants. Contrast that with just 0.55% of Shanes. Arnolds therefore appear to have a much higher tendency to be accountants than Shanes
implies they're listing the top P(J|N) values for each J. *(edit: they're comparing Arnolds to Shanes, not Arnolds to all accountants?) I think your approach is the most consistent but is it what they're using?
As I read it, it means that, say, 1% of Elwoods become farmers, while only .1% of Steves are farmers. That is, if your name is Elwood, you are much more likely to be a farmer than if your name is anything else.
http://en.wikipedia.org/wiki/Nominative_determinism
Think someone called Crapper being a toilet salesman, or Smith being a Blacksmith, Baker being a Baker etc.
I.e. a funny non-science based upon misunderstanding historical coincidence (profession being used as surnames, family marketing name being used as a crass synoniem for something else) with causation.
https://www.google.com.au/webhp?sourceid=chrome-instant&ion=...
If 99% of farmers are Elwoods, you can't claim that one's name being Elwood means one is more likely to become a farmer.
If anything, it's more like an ecological fallacy or unwarranted extrapolation to the future.
Apparently, if your name is Elwood you're more likely to be a farmer today. But that obviously doesn't mean that kids named Elwood are more likely to become farmers (it could, but I'm afraid you're gonna have to proof that).
Wikipedia has some simple examples: https://en.wikipedia.org/wiki/Correlation_does_not_imply_cau...
I'm sorry I can't provide sources.
http://freakonomics.com/2009/04/24/yes-part-ii/
referencing this paper "Why Susie Sells Seashells by the Seashore: Implicit Egotism and Life Decisions" http://www.stat.columbia.edu/~gelman/stuff_for_blog/susie.pd...
[1]http://www.amazon.com/Yes-Scientifically-Proven-Ways-Persuas...
These days, you pretty much have to put a pretty picture in an article if you want traffic from social sites.
I mean, the information was interesting enough to me to try to read it, but I frankly can not parse what looks like #FFFFFF text on #FEFFFE background.
For the love of god, If you want to present data like this, ensure at least a little contrast exists between the text and background
Direct link to the blog that goes in to a few details of how they work out the names.
Nearly every one was John Gallant, the surname is very common and it's a small community.
The funny thing is the surname is so common people have nicknames such as John 'Rabbit' Gallant but that name becomes so well-known his son will be called Rabbit Jr.
So the plaque is full of a wild mix of actual surnames and given names and of course nicknames but also the junior of the nicknamed people.
Add to that one family has seven daughters all named Mary.
Songwriter is interesting too: 4 out of 5 end with a variation of "y".
Their birth certificates could very well read Daniel, William, Michael, James, Richard, Steven.
Similarly, the EE's probably have a tendency to be more formal on their resumes or business cards.
Their family and buddies probably calls them Bernie, Gene, Eddie, Chuck, Freddy, Harv.
Chances are that there are largely unsurprising gender/age/class/race biases in the full dataset too, but these have been selected for effect (I bet most of the rest of the disproportionately common football coach names are pretty regular names for men born between 30 and 55 years ago with some of them possibly even having additional syllables, and similarly and wouldn't be surprised to find that Jim and Bill, for example, also featured high in the list of people disproportionately likely to be electrical engineers, possibly even above Eugene)
For example: http://benedictcumberbatchgenerator.tumblr.com/
edit: (or yeah, what MichaelTieso said :) )
There are lots of common surnames that demonstrate this - Taylor, Cooper, Cobbler, Smith, and so on. Why not given names?
Note that this was pre legalisation so watching peaky blinders(uk version of Atlantic boardwalk) on the BBC was interesting shall we say
http://andrewgelman.com/2005/08/05/dennis_the_denv/ - short summary.
You don't say... :) Perhaps because names starting with K are simply very common in Poland? ;) 1 in 6 out of 100 most popular last names starts with K.
And quite more IT employees than whom exactly? Presidents of the country? Since 1989 Poland had: Jaruzelski, Walesa, Kwasniewski, Kaczynski and Komorowski... Among prime ministers (since the 90s) about 1 in 5 had either first name or last name that begins with K :)
Polish article: http://wiadomosci.gazeta.pl/wiadomosci/1,114873,9855837,Krak...
By the way, believe it or not, I do work as a programmer (in Poland), and my first name does indeed start with K :) Not the last name though, so I feel safe.
While the serial killer theory could boil down to a statistical oddity, it reminded me of this novel by Lem: http://en.wikipedia.org/wiki/The_Investigation
It doesn't say they're the top 6, just that they're "6 of" -- and having worked with a lot of similar data sets in the past, the results here feel a little overly edited (i.e. exaggerated, stereotyped) to me. I'd be happy to be proven wrong, though.
There they list the actual top 5 for some professions. Having read that, I am inclined to agree with you here.
The top 6 for car salesmen in that graph has literally only one name (Clay) that is in the actual top 5 and even that was only 5th. The top 4 names (Emmett, Luther, Emanuel, Morton) all got replaced with stereotypical white working class guy names.
The top 6 for surgeon in the graph has no female names yet the actual most disproportionately common name for surgeons is 'Vivienne'.
Not that I don't love both.
http://fivethirtyeight.com/features/how-to-tell-someones-age...
My father is a Pediatrician, and he has always commented on the strong correlations between relatively common names in his current crop of patients and the names of 5-year-ago-popular TV shows' protagonists.
The first time I was able to make the connection, it was Brandon/Brenda/Dylan.
Once the stats were adjusted for the above, I suppose we'd end up with not much more but noise and some spurious correlations occassionally: http://twentytwowords.com/funny-graphs-show-correlation-betw...
Seems sort of ... logical.
Also, WHY are things separated by color and opacity? If we're going by opacity, apparently one of the most populous and important professions is...race car driver. Really? That's one of the few professions on that huge infographic that is at maximum opacity?
http://www.verdantlabs.com/blog/2014/12/30/names-by-professi...
You see where this is going. If you correlate the data with the popularity of names in general, you'll find that Arnold is a much more popular name than Shane ...
I think I am, but I might be wrong on that one too.
Given a pool of 10 people to hire from where 6 people are named Arnold and 4 are named shawn Shawn, wouldn't you expect that same relation to show in a specific profession? You can, of course, go ahead and compare those relations. So, if there only were 5 race car drivers in the world and 4 of them were named Arnold that would be noteworthy (as opposed to 60% of them, which is what you'd expect). Please let me know if I got something wrong.
---
Somebody just posted an article stating the numbers where actually correlated with the frequency they are used. I'm not sure how they did it though given that those numbers change over time.
You would expect that the percent of each profession by name should be the same. So if there are 25 professions, then within any given name you should have 4% going to each professions. So 4% of Arnolds should be a profession, and 4% of Shawns should also have the same profession. That is not what the data shows though. Within any given name there is a tendency towards the professions the chart shows.
This definitely does not account for location, birthdays, etc. But it is still interesting.
Also, many cultures obsess over giving the child a good name with a good meaning as they think it determines their future.
Aha funny how people don't define their data correctly, this is an important datapoint
People of a certain social class are more likely to name their child certain names, and those children are more likely to grow up into particular professions.
Name inequality represents clear evidence against the existence of a meritocracy.
Meritocracy in what sense?
Coming from wealthier background, you're more likely to be well educated.
While we may dislike that it is so, it doesn't mean there's no meritocracy in the sense that employers don't hire people based on their qualifications alone. Only that they don't care where these qualifications come from.
So, by your admission, education is not meritocratic.
I claim that employment is non-meritocratic first by your measure: if access to a better education is not merited, then employers concentrating on qualifications alone are not hiring according to merit.
I claim also that employment is non-meritocratic independently. The most obvious example is that wealth similarly gives you exclusive access to low-paying but prestigious jobs.
I suspect overall that background wealth is still a better marker for job status than education.
That doesn't follow. Maybe education changes your merit, and people with a better education actually are better at their jobs.
Or, in your meritocracy, people pay to increase their merit/the merit of their children? That contradicts my definition.
Then you might be able to say that everyone had access to a decent standard of education.
We don't know if genetics is an effect here because we can't eliminate background wealth - even twin studies are broken because twins get adopted into similar environments.
But, the suggestion from twin studies is that income is primarily environmental, i.e. not genetic. Here's a fresh reference: http://www.sciencedaily.com/releases/2014/11/141106113202.ht...
I suggest that it's fashionable to say social inequality is down to genetics - it certainly happens a lot in these forums and it comes up especially when people would rather pass the buck for difficult social problems.
When people on HN talk about the meritocracy as though it is something that exists, they are talking about it in the tech industry, not in the education system that precedes it.
Putting the unfair part in an early funnelling step of the hiring process, and then abstracting that away, doesn't change that the whole process is unfair.