Note how the article mentions that the gender can be determined after a few keystrokes, even though the user never entered that specific information. This is certainly not the only metric that can be identified. The point of the article is to develop a solution to prevent leakage of private/personal information.
Research got median 88% accuracy testing subsets of 98 males and 35 females.
Note that I got 74% accuracy on that data set by guessing male, male, male, male, male...
( You have a very ironic username given the circumstances. )
In general, I don't believe it is possible to distinguish male and female typing patterns.
What you might be recognising is how people learned to type combined with the size of their hands - that might partly but not exactly break along gender lines. Bucketing people on that basis is just a recipe for awkwardness.
Quote from the paper: We use the public GREYC keystroke benchmark database for this work. It is one of the largest databases (in term of number of users and sessions) in keystroke dynamics. To out knowledge, no existing database contains more individuals. In order to reduce the bias due to this high quantity of male information, we only kept the first n male samples( where n is the number of female samples).
( Don't bother with your response, I won't be reading it. )
Yes. That's their own database which they're talking up, the one that they made to do this research. That's what I was talking about.
>In order to reduce the bias due to this high quantity of male information, we only kept the first n male samples( where n is the number of female samples).
It happens that I didn't read this part.
On reflection, what I understand now is far worse than what I originally understood:
- They have 35 females and 98 males, they take many handwriting samples from each.
- Since the participants provided many samples, these samples appear both in the training set data and in the test set data.
- I use the training set data to figure out if I can recognise the handwriting of the 35 female participants.
- Then I look through the test data to see if I can identify those participants again.
Basically what you've shown is you can identify the handwriting of 35 people if you've already seen it - 88% of the time.
Splitting groups into 'female' and 'male' is a red herring. This method would presumably work, even if I split them into two random groups.
If I'm right, this is not even state-of-the-art. In 2006 they could have been scoring 96%: http://abcnews.go.com/Technology/story?id=97978&page=2
How would mobile work with this? Sometimes I use both fingers, sometimes just my thumb on one hand. Would it just create multiple behavior profiles for me that are accepted?
In combination of the right password and the behavior match, it seems like this would actually be pretty strong. I'm looking forward to trying to break it tomorrow with a friend.
You could use it as an indicator and trigger a warning email.