NeuralTalk2: Efficient Image Captioning code in Torch, runs on GPU
github.com
github.com
Something people don't fully appreciate about neural networks is that their performance is quite a strong function of their training data. In this case the training data is taken from the MS COCO dataset (http://mscoco.org/explore/). That's why, for example, when Kyle points the camera at himself the model says something along the lines of "man with a suit and tie" - there is a very strong correlation between that kind of an image in the data, and the presence of a suit and tie. With such a strong correlation the model doesn't have a chance to tease the two concepts apart. A similar problem would come up with an ImageNet model, where a similar image might be classified as "seatbelt", because there is no Person class there, and shots of people in that pose usually come from the seatbelt class. It happens to be the most similar concept in the data it has seen. Another example is if you pointed the model at trees it might hallucinate a giraffe, since the two are strongly correlated in the data. Or when Kyle points the camera at the ground I'm fully expecting it to say relatively random things, because I know that those kinds of images are very rare in the training data.
In other words, a lot of the "mistakes" are limitations of training data and its variety rather than something to do with the model itself, and it's easier to recognize this if you're familiar with the training data and its classes and distribution.
One simple low-hanging-fruit approach would be to include a large repository of additional data (e.g. all of ImageNet) and label it all as a "garbage" class. This way the model could at least learn to distinguish the kinds of images in its training data from the universe of images, and this could be used as one proxy of confidence.
Another simple proxy is to look at the probability of the generated sample, since usually the model tends to assign more diffuse probabilities in more uncertain cases. But this is also not a very clean approach for various reasons.
Another, and probably most appealing, approach would be something along the lines of Bayesian Neural Networks, ensembles, or approximations with dropout, where the disagreement between the predictions of all submodels can be used.
One more idea would be to have a (non-differentiable / REINFORCE?) penalty based on sentence likelihood using doc2vec or skip-thoughts to avoid the "blah and a blah on a blah" type errors that seem to be common in captioning.
One more would be to use TV / YouTube captions, but that data is extremely noisy - even more the COCO captions, unfortunately!
It feels like getting something for free :) Of course there is a limit to how much signal you can extract from a noisy dataset, but the amount of time and human energy invested into creating and improving datasets can be quite large relative to finding another cool trick that can improve performance.
However, I wonder which will come first to make these systems "robust" for the average joe's real world uses for these perceptual systems; a large, well labeled dataset or more transformations and semi-supervised learning approaches?
[1] http://www.cs.stanford.edu/~acoates/papers/coatesleeng_aista... [2] http://cseweb.ucsd.edu/~elkan/posonly.pdf
If so, I'm guessing Google, maybe Facebook too, has plenty of data. What else is holding them back?
Also it's not only the size of the dataset, it's also the size/variety in the label space. ImageNet is quite comprehensive, with many varied labels. MS COCO is quite biased towards a narrow ~hundred classes.
I'd love to see a properly large dataset of images "from the wild", with no restrictions on content (unlike what is done in MS COCO), annotated with sentences. From my experience with adding data to models in these situations I'm quite certain this would work _significantly_ better.
Google's Show and Tell seems considerably superior to competing approaches.
What's different between the two is the engineering portion: NeuralTalk was written in Python and ran on CPU without batches (so probably it did not converge as well), it did not do CNN finetuning (which helps a ton), and it did not do ensembles, and I used VGG while Oriol used GoogLeNet (though I don't expect much difference there). There are a few more tricks that Oriol talks about that give you a few small extra points, but that's basically it. I also don't want to take away from Oriol's top result (these small tricks and engineering are very valuable and hard to come up with), but I would be hesitant to draw conclusions about which models work best by looking at the model diagrams in a Figure, and comparing to numbers in the table.
That's why I've over time grown cynical about results in tables of papers that compare one work to another - there are too many variables to keep track of and it's confusing unless you know the full details. The truth is that results in tables are model + engineering + noise. The papers pretend that it's all model, but in fact the latter 2 have a huge impact. At least results that only compare two models in the same framework can be trusted a bit more to contain information, but even then you find that in practice people can be a lot more generous to their own models than their baselines. Hence the famous saying: "The second best model in the paper is in fact what you want to use".
To answer the original question though: The idea presented in Show Attend and Tell (spatial attention during caption generation) is clearly good and if compared properly I'm confident would turn out to work better. Also, the Berkeley paper seemed to have a nice architecture where there was a 2-layer LSTM but the image was only plugged in on the second layer (1st layer was an image-independent Language Model). In their controlled experiments this seemed to work well, I think it's an interesting idea, and I'd want to try to reproduce it. The models presented in my work and Oriol's are the simplest architecture that has the core nugget of the approach and that gets the job done, but I'd expect many bells and whistles on top of this to work better.
That's why I like Kaggle competitions: most approaches tried on a given challenge will likely be optimized to their very limits, so that it's the nature of the different approaches that ends up making the difference.
This algorithm needs badly temporal dimension, some kind of short term memory that lets it interpret using context. At the very least to filter out freaky readings of a train station when looking at the ground, best case scenario it would enable building deeper understating of its surroundings. Maybe not even memory, but Bayes filter to prime next estimation. Then throw movies at it.
Even as it is this could be adapted for the blind. I can imagine app that will simply build a model of what it sees and answer questions or warn about stairs/walls/roads/other dangers. There isnt all that much to make it as clever as a guide dog.
Do you plan to add beamsearch?