Using Machine Learning and Node.js to detect the gender of Instagram Users
totems.co
totems.co
There is also no evidence of doing cross-validation, and in another comment they say they used entire data set to do variable selection - a pretty bad mistake. They justify by saying they aren't in an academic environment, but thats kind of a bad excuse, as given the way they've done it I'm very unsure whether they actually are getting the accuracy they think they are.
I also worry that they sunk two man-months into this when they could probably have achieved similar if not better results with off-the-shelf and battled-tested tools. Sets off a lot of warning bells.
Although their computing is very smart, the output cant be better than the input they used. Just determining the gender by looking up a related Facebook-Profile should therefore be a better solution in my opinion.
Merely choosing to withhold information about yourself does not insulate you from a breach of privacy. That others do disclose such information allows 3rd parties to make really good guesses and inferences about you.
There's a strange morality here: at what point is it unethical to voluntarily disclose data about oneself, if it could be used in a way to harm someone else's privacy? Short of drawing a moral boundary (it could very well be impossible), we might do well to at least acknowledge the cost to these methods, alongside their benefits.
That's an interesting question. Especially since the data you disclose may triggers inappropriate inference of characteristics on someone else, maybe eventually causing some form of harm (anytime the demo fails to classify someone, we do cause some harm to him/her in a way). In the case where the misclassification is more harmful than the privacy disclosure, one is better off disclosing the information... weird equilibrium.
At what point is it unethical to exhale, given that carbon dioxide is toxic to humans and is a greenhouse gas? At what point is it unethical to vote, given that you might influence an election in a way that is bad for society or some subset of society?
It's true that nearly every (and the "nearly" is just a hedge) action we take has some negative externality. I personally don't lose sleep over the ones that are virtually impossible to measure.
Mildly alarmed to learn I'm only .039 probability male, though - better bloke it up on Instagram.
That said, I did get 0.998 female and 0.996 male. Oh well.
Why implement the training in NodeJS and not use an existing library in R or Python (scikit-learn) and just implement the scoring (feedforward network) in Node?
Did you just use a single test/train split? What is the variation in Res if you run cross validation?
Your article suggests that you used MI to select the 10k best features. Did you perform this MI feature selection before your test/train split? If so, you would already be "using" your class labels, and the results will be biased. It is likely your true generalisation error will be lower.
We wanted to contribute to the nodeJS ecosystem and build whatever tool was missing to use neural network directly from NodeJS or at least as an add-on. We also wanted to come up with a simple an straightforward implementation to serve as an educational example rather than just bind into an existing library (even though the results might have been better of course)
> Did you just use a single test/train split? What is the variation in Res if you run cross validation?
We didn't use cross-validation but rather simple train/test split (though our test set was quite large ~100k / 570k). As explained in the intro we wanted to stay very practical and were ok with dirty shortcuts as long as the result looked OK.
> Did you perform this MI feature selection before your test/train split? If so, you would already be "using" your class labels, and the results will be biased. It is likely your true generalisation error will be lower.
Yes MI selection was made on the overall data set before training. You totally are right that this is a bias against the test set. Nice catch.
double dW = alpha_ * val_[l][j] * D_[l+1][i] + beta_ * dW_[l+1][i][j];
W_[l+1][i][j] += dW;
If you want to get an output class probability, softmax is the standard way. Minimize KL-divergence instead of squared error.You don't seem to be doing any regularization. It could maybe give you better generalization.
I think you could get a speedup by doing your linalg with blas, I guess this would complicate the code though, making it a tradeof.
Training on multiple threads and averaging is a nice touch. It would be interesting to hear if (how much) it improved your results.
I think we used what is described in Artificial Intelligence: A Modern Approach... But I have to check because what you propose seems better.
> If you want to get an output class probability, softmax is the standard way. Minimize KL-divergence instead of squared error.
Thanks! We'll totally try that.
> You don't seem to be doing any regularization. It could maybe give you better generalization.
Thanks again. Someone mentioned that before as well. We'll have to experiment with that as well.
> Training on multiple threads and averaging is a nice touch. It would be interesting to hear if (how much) it improved your results.
Training was much faster and therefore tractable on a much larger set but we didn't manage to get our best results using this multi-threaded approach unfortunately as described in the post.
Maybe with a bigger training set we could have reach better results using multi-threaded training. That being said, the averaging phase disrupts a lot the overall backpropagation process, so I don't know how efficient it can be... Some advanced experimentation would probably be interesting here.
oops maybe I spoke too soon, allow me to backpedal a little. I still recommend minimizing KL-divergence.
What seems odd is that the "test tool" allows you to tweet whether it's wrong or right. Why not just have it make a call to your API or something to tell you directly, so you can look at the profiles and figure out what's gone wrong?
Thank you for sharing the C version, I'll use it for sure.
I know it's not perfect... But heh. Hope it's ok.
My instagram name is the same as my HN user id, and you classified me with 99.3% as female... Needs work!
The algorithm could probably be improved to also take the instagram name into consideration. Someone named "arthur" is very unlikely to be female.
Interesting, since Instagram's API only allows 5,000 requests per hour, (http://instagram.com/developer/limits/) and does not support bulk requests of user data. How does this application bypass this limit?
Actually, Instagram API limit is pretty high when compared to other platform. Today we have something like 100k tokens available to us, which means we can make 12bn+ calls everyday. Almost like having a firehose. We don't use all of them but we're one of the top users. Though there are at least 10 bigger users than us on the API (according to them of course). Hope it helps!
I'm fairly certain this isn't an intended use of the API. Kudos to you for posting your method here, but I wouldn't be surprised if your app gets banned because of this (a lot of people at Instagram read HN).
For those curious, IIRC it took a while to get our account's limits raised, and we had to implement some request caching to stay under the limits as much as possible. All around, it was interesting.
Yes, for using what are intended as per-user activity tokens for public scraping (which the user who has been issued the token has not requested). As you said, you can assemble a firehose using this method, and if they'd wanted apps to access a firehose, they'd have come up with an API for it.
This is a quite idealistic view of the problem. But it probably holds some truth I have to admit.
That being said... studying bayesian networks more thoroughly might raise better results indeed. Don't know though if Gmail is using bayesian networks or deep learning?
In the few cases where NN have achieved new state of the art on text, the stanford sentiment analysis work and a few more recent works, a full sentence parse is needed. Sentence parses do achieve 95% accuracy, but only on well structured text in a given domain. Plus, they are hugely time intensive compared to the large scale linear classifiers like vowpal wabbit or sophia-ml.
Regarding perceptrons, a basic perceptron, although it is a linear classifier, will not achieve state of the art. Averaged perceptrons get you closer, but what you really want is a discriminatively trained linear classifier with regularization.
If I had to bet, gmail is probably using something closer to https://code.google.com/p/sofia-ml/ than a NN. Maybe a Googler will surprise me though!
Yup, the Google "Priority Inbox" feature does indeed use a linear classifier, in particular logistic regression [1] for the reason of scale as you point out.
Also, IIRC Gmail's original spam detection used naive bayes. It may have evolved since then.
[1] http://static.googleusercontent.com/media/research.google.co...
You can think of naive bayes and a perceptron as roughly equivalent in terms of expressiveness--they're both linear models--but a perceptron is usually better since it can account for correlations between input variables.
As you say, a perceptron is a one-layer neural network, so with a large enough training set, a multi-layer neural network will almost certainly perform better since it can recognize combinations of features that work well together.
Bayesian filtering for spam detection is a good starting point since it's easy to implement, and was very popular in the mid-2000s, but with all the advancements in deep learning since 2006, I'd almost certainly bet on a neural network these days.
PROBABILITY FEMALE: 0.997
PROBABILITY MALE: 0.569
I wonder if the fact that I mostly just post pictures with no text accompanying them skews things. PROBABILITY FEMALE: 0.003
PROBABILITY MALE: 0.001However, my business has a 0.885 probability of being a woman, which is odd for a men's brand.
Errr, so it's out of 1.002?
PROBABILITY FEMALE: 0.003
PROBABILITY MALE: 0.996
I would say this doesn't work very well.
Also, Tegan and Sara are great singers and artists, but neither of them is an exemplification of what our culture considers stereotypically female.