Show and Tell: Image captioning open sourced in TensorFlow
research.googleblog.com
research.googleblog.com
Google clearly has many different trained versions of this network sitting around. It must have been a conscious decision not to release them. Is the point to artificially create a barrier to entry for hobbyists that might want to apply this research? If so, why bother releasing it at all? I'm really scratching my head here.
In order to get full AI we need to add behavior and embodiment to these meaning vectors. They need to be trained by reinforcement learning, to learn the behavior that maximizes rewards. Meaning vectors are just a small part of the final system, equivalent to our ability to see and speak. The most difficult part is that of learning behavior.
We seem to have more or less solved perception. Given a high dimensional, "raw" input space, we know how to process it into a more usable, more abstract representation.
Work on perception has also given us certain limited kinds of behavior. We can generate images from abstract representations by inverting our image recognition architectures. RNNs can do perception over sequences (e.g. of words), but can also generate language output or a sequence of control commands for robotics.
One area in which I think human cognition is far ahead of neural network research is control flow. Maybe there's a better name for this. We seem to be able to encounter a new cognitive challenge and quickly design a mental program to solve it. We can attend to relevant sensory streams, various kinds of memory, and then design (sometimes novel) behavior to solve the problem.
Work on attention and architectures with more sophisticated internal representations like stack-augmented RNNs are definitely moving in this direction, but it seems like we have much further to go on this front than in visual perception, for example.
It's one of the papers that kicked off this approach to image captioning, but is much more ambitious. Lots of things in it don't really quite work, but it shows where this work is going.
Train it from github.com's commits, logs - auto learn "what the SW bugs look like" and scan for new one....
Automatic Patch Generation by Learning Correct Code
TensorFlow is just a framework (as are Theano, Torch or DL4J) for expressing the network architecture. Framework:Network ~ ProgrammingLanguage:Algorithm
Running it locally on the user's machine would take far too long to train, especially as you would have to use the CPU in the majority of cases since many people don't have a separate GPU.
Training is another matter.
> The time required to train the Show and Tell model depends on your specific hardware and computational capacity. > In this guide we assume you will be running training on a single machine with a GPU. In our experience on an NVIDIA Tesla K20m GPU the initial training phase takes 1-2 weeks. The second training phase may take several additional weeks to achieve peak performance (but you can stop this phase early and still get reasonable results).
> It is possible to achieve a speed-up by implementing distributed training across a cluster of machines with GPUs, > but that is not covered in this guide.
> Whilst it is possible to run this code on a CPU, beware that this may be approximately 10 times slower.
So, I assume it will take veeery long time to train it on a MBP, unless they publish their pre-trained data-set.
[1] https://github.com/tensorflow/models/tree/master/im2txt#a-no...