How Quid uses deep learning with small data
quid.com
quid.com
But I'm wondering how you get around that with the neural net. In the post, you said there are only a few hundred labeled examples, right? How can a neural net with hundreds of parameters set those parameters to anything reasonable, and not overfit, when there are about as many parameters as examples?
Edit: Also, do you think more data or a "better" or "more sophisticated" model would make the results better? I would guess more data would trump better model, but not sure.
Yes, I'd definitely agree- more data is what we need here for further model improvements.
Basically I have simpler interfaces and the ability for multiple people to quickly answer questions like this on a set of data. Easily exportable in the end as well. If you're interested in using that to get some more data on sentences, let me know. I'm really curious how much better the results get with more data, and this could help.
[0] https://bigishdata.com/2016/11/01/classifying-country-music-...
So the question is: does the OP really show good generalisation?
It's hard to see how one would even begin to test this, in the case of the OP. The OP describes an experiment where a few hundred instances were drawn from a set of 50K, and used both for training and testing (by holding out a few, rather than cross-validating, if I got that right).
I guess one way to go about it is to use the trained model to label your unseen data (the rest of the 50k) and then go through that model-labelled data by hand, and try to figure out how well the model did.
We're talking here about natural language, however, where the domain is so vast that even the full 50k instances are very few to learn well. That doesn't have to do anything with the model being trained, deep or shallow. It has everything to do with the fact that you can say the same thing in 100k different ways, and still not exhaust all the ways to say that one thing. So 50k examples are either not enough examples of different ways to say the same thing, or not enough examples of the different things you can say, or, most probably, both.
It's also worth remembering that deep nets can overfit much worse than other methods, exactly because they are so good at memorising training data. It's very hard to figure out what a deep net is really learning, but it would not be at all surprising to find out that your "powerful" model is just a very expensive alternative to Ctrl + C.
It's just memorised your examples, see?
I do neuroscience research and where I am coming from I have maybe n=150to200 per class. And that is not generally regarded as a tiny sample.
n=50000 of text data is different, since there will be less repetition of contextual structures and words (particularly with proper nouns). The fact that the dataset only uses "hundreds" as mentioned in the original post is interesting.
Actually Kim's model you're using doesn't require padding because it uses k-Max over time pooling.
Also kuddos for NOT updating your word embeddings during training! A lot of people are doing it, but IMHO it's a mistake most of the time.
Can I have that for my email? (Seriously) And as browser plugin? Oh and on telephone, TV, radio and in real-life would be also nice.
It's probably also a nice predictor of startup success, developer quality and sales guy effectiveness.
I just wonder if I would ever read or hear a Politician again.
Very inspiring...
I thought the whole point of "deep learning" is its approach to using data.
However, picking millions of parameters with small amounts of data is unlikely to work well.
This is a normative statement, do you have empirical evidence?