Determining Gender of a Name with 80% Accuracy Using Only Three Features
blog.ayoungprogrammer.com
blog.ayoungprogrammer.com
(Regarding sex vs. gender: yes, they aren't perfectly correlated and some people don't stay with their assigned gender, but AFAIK they often/usually choose a new name which probably matches their gender? Thus why "dead names" are a thing.)
Here's one example I've created regarding the Pulitzer board composition: https://github.com/compciv/gendered-pulitzer-board
Note: it was only an example...obviously a little tweaking is needed to put into actual production. Also, obviously has varied effectiveness on datasets with non-traditional American names. Here's one of the more comprehensive efforts by a student, on New Yorker bylines:
https://github.com/alecglassford/compciv-2016/blob/master/pr...
The SSN name data seems almost certainly flawed to a small degree...I.e. It's just hard to believe that there are dozens of boys named Jennifer (and yet, strangely, no boys named Sue!)...but we're talking about infinitesimal rounding errors. The vast, vast majority of names are 99% one way or the other...with a few exceptions such as Leslie...though you can mitigate that by using older years of the SSN database.
And here's a battery of SQL queries relating to baby names and gender analysis: http://2015.padjo.org/tutorials/babynames-and-college-salari...
So I have to strongly disagree with OP that 80% accuracy is something to be astounded by when it comes to gender classification...however, I do agree that in terms of features, doing a simple frequency count of last characters...or number of vowels overall can be a strong indicator of gender for a name, with female names tending for the softer sound. I wonder how much more using Soundex would add to the accuracy? Creating a trained name classifier would be a fun project in service of a tool that could gender classify how masculine or feminine a made-up name sounds like...which would be a slightly useful tool if you were a fantasy fiction writer, though I suppose if you were to be a successful writer, your ear would be trained well enough for he purpose to not delegate it to a computational tool.
I agree that the SSN data has flaws, but I only took names with at least 20 people. But the classification is probably iffy, as some names are classified as male and female.
> So I have to strongly disagree with OP that 80% accuracy is something to be astounded by when it comes to gender classification...
I originally hypothesized, I could reach 90% accuracy, but I could only get up to 82% max. As stated in the blog, 80% is the accuracy of a mammogram detecting cancer in a 40-45 year old woman which is pretty good for 3 features!
> I wonder how much more using Soundex would add to the accuracy? Creating a trained name classifier would be a fun project in service of a tool that could gender classify how masculine or feminine a made-up name sounds like...which would be a slightly useful tool if you were a fantasy fiction writer, though I suppose if you were to be a successful writer, your ear would be trained well enough for he purpose to not delegate it to a computational tool.
This would be very interesting to see!
I used to work with Text-Mining techniques and it was quite easy to come up with a statistical metric that could reach 70%-75% accuracy, for almost all problems that I've dealt with (mainly in the sub-area of concept extraction - information retrieval). I assume it is similar for this guy, so that is why he reaches this kind of values very fast. For instance, with only the character frequency and order he reports results of 75% which, if you look carefully, is just 20%-25% above the random baseline, not "pure" 75%..
Then he trains a classifier (Random Forest) and extracts the most relevant features that the classifier used, which is existence of an 'a' character in the last two positions of the name (the order of 'a' feature seems quite irrelevant to me). With these features he reaches the 80% ("random" + 30%) accuracy which is a good number but still far. I would say that these two features are to be expected, as, at least in my mother tongue (Portuguese), feminine names end with an 'a' (eg. Maria, Ana, Roberta [vs Roberto - masc], etc..).
But all in all it can be considered a good exercise!
Still impressive but it would be interesting to see the result for other countries. Maybe for some countries, the gender could be harder or easier to guess. I can see how some people would use this to try to reconstruct a database where gender is missing.
Counterexamples among the 50 most common male names given in 2014 are:
Andrea 2.46%
Mattia 2.39%
Luca 1.14%
Elia 0.45%
Counterexamples among the 50 most common female names given in 2014 are: None
There are a handful of names ending with E or a consonant for both genders, which will have to be placed at random.Source (in Italian): http://www.istat.it/it/prodotti/contenuti-interattivi/calcol...
Same in polish. If 'a' is last you get like 98% probability it's female name and if last is not 'a' then 98% it's a male name.
Thus many Romance countries (and Christian countries, since many The Catholic Church requires the child to take the name of a saint for Baptism), have a disproportionate number of female names that end in "a."
The "a" phenom signifying femininity is even suspected to go all the way back to Proto-Indo-European.
Obviously as an exercise in machine learning, this is perfectly acceptable, but as a solution to a problem, not so much.
It is very likely that the name absent from the db will have features that reveal its gender.
So 1) try db lookup, if absent 2) use a classification algorithm (which will always provide an answer)
This is definitely some interesting work, but I really would recommend not using it in the real world. You almost certainly don't need to know a user's gender; if you really do then you should ask them what it is and use their own definition.
I agree with your above points about the dangers of making guesses on gender. I'm just not well versed in the progressive gender concepts and don't understand the semantics you are arguing for.
I appreciate the point you're trying to make, but it's kind of a straw man argument.
While I'm not really buying your root argument, this statement is spot on.
Case in point - I'm a man named Lyndsy.
You'd have to be speaking a language with gender marked on the noun; English is not such a language.
actor/actress
fiancé/fiancée
widower/widow
waiter/waitress
in modern American English, "who" is a pronoun for people (and anything being granted a sort of honorary personhood, like a pet), and "what" is a pronoun for non-people. The third-person pronouns follow the same distinction, with "it" being marked for nonpersonhood. This is a change from historical usage, and is why e.g. some people today will be offended by another person referring to their baby as "it".
i) you want to know how a person identifies
ii) you're gathering demographic data so you also want to get trans status
At this point there's no point asking about gender. Certainly in the UK asking for sex and then asking for gender is going to be seen as hostile to trans people, while asking for what sex someone identifies as and then asking about trans status is seen as less hostile.
I can't think of a situation where you'd ask for the sex and not for trans status.
Honestly, I don't think you need to ask for any (gender|sex)-information in most cases, but it is often added out of completism. The only valid reasons are: - Self-presentation of the user, e.g. in a profile page -> allow the user to write whatever they want - Advertising, argh, but you often can't get around this. I don't know if advertisers care about how people identify or about trans people, but they probably want a simple m/f flag that fits enough people. - Medical or legal stuff
Seems that gender can be used both ways.
Somewhat orthogonal is the issue that the guess is binary (m/f). But the thing is, as long as the accurracy is ~80%, the total mismatch rate is much larger than the rate of people that are mismatched because there is no category for them.
Obviously, you shouldn't guess someone's sex and or gender and present it to them or others. But this can be used for a lot of other things:
- If you have a large dataset of names and other information, but no gender info, and want to see if e.g. women are underrepresented. Even with a weak classifier, you can set bounds.
- If you have a real-name, real-data policy, you can use this as a pre-screening to see if the information entered is plausible. It's of course up to you to then do something responsible with that information. I'd prefer to allow pseudonyms or assumed identites in most cases, but sometimes it's not possible (e.g. if this is a egovernment or insurance project).
If people want to lie, they are going to give you a fake gender to match their fake name. What you're suggesting would flag 20% of your legitimate users and 0% of malicious users.
Unless you are specifically targeting liberal arts college students or Tumblr's otherkin community, Gender: Male/Female is going to cover the vast majority of your users. Unless you are running an adult dating website, there's no reason to ask your users' biological sex ever.
All that being said, I do agree that if you want to know a user's gender you should be asking them rather than trying to guess. And since it doesn't cost anything extra you might as well throw "other" in there as an option.
You simply find the most common names in each country with predominant language. Want to find Indian names? Most common names in India. Latino names? Most common names in Mexico. etc. It's not perfect, but nothing is. Case in point:
As a wise man once sang, "I'm not black like Barry White. No, I am white like Frank Black is."
Maybe this is a failure of imagination on my part though... Do other people have a sense of some altruistic feature that would rely on tech like this?
Even if the computer improves to twice human accuracy (i.e. mistakes half as often) I don't see how it'll reveal or bias anything that wasn't already about to happen.
I haven't seen Icelanders rising up to make their names more ambiguous, and Italians don't seem unduly overrun with nefarious lipstick ads.
I also guarantee that advertisers can do a good deal better than 80% by browsing habits alone.
I'm writing to you to let you know that as you haven't provided us with a preferred form of address, out computer has automatically chosen one for you (see above). If you'd prefer us to use a different form, please let us know.
Then it is your imagination being nefarious.