Attention Is All You Need
papers.nips.cc
papers.nips.cc
Its a great time to be alive!
(The next step for me would be to follow the citation trail to the original paper, but that might not be the best place to come to an understanding of the thing.)
An imperfect analogy is how the human visual system has better resolution at your eye-line's center than at it's edges. In this analogy, your brain should not waste effort processing image details in your peripheral vision.
DNNs typically operate on fixed-size tensors (often with a variable batch size, which you can safely ignore). In order to incorporate a non-fixed size tensor, you need some way of converting it into a fixed size. For example, processing a sentence of variable length into a single prediction value. You have many choices for methods of combining the tensors from each token in the sentence - max, min, mean, median, sum, etc etc. Attention is a weighted mean, where the weights are computed based on a query, key, and value. The query might represent something you know about the sentence or the context (“this is a sentence from a toaster review”), the key represents something you know about each token (“this is the word embedding tensor”), and the value is the tensor you want to use for the weighted mean.
Still... even though I understand what works and what doesn't, I don't really understand why RNNs are better at processing symbols and CNNs are better at processing images. They hold the same information, and work the same way, they are just organized differently. It makes little sense.
ox = w1 * x + w3 * y + bx
oy = w2 * x + w4 * y + by
And then learns how to change w1 ... w4, bx and by to get ox and oy to be more useful (less error). Is it so weird that a 2 layer network (which is the same formula, but it uses ox and oy as inputs to a second identical calculation to produce the final outputs) will produce better results, even with learning ?
Even in a neural network so absurdly large as the human mind, it isn't just a bunch of wires connecting everything to everything. There are clear patterns, clear formulas that are dictated not by learning, but by the architecture itself (e.g. CNNs in the eyes and optical cortex, and something much more like LSTMs in memory heavy regions).
In some sense one might even say it doesn't make a difference. A fully connected network with the same fanout as a CNN can do everything a CNN can, and more. Likewise, a network that is simply presented with the last 50 timesteps can do strictly better than an LSTM or GRU RNN would.
But such a network would have much, much greater computational complexity (by a factor in the thousands at least, in the case of those CNNs by a factor that is roughly the number of pixels in the image).
So they don't work better per se. In fact they are "worse" by the most important metric (error). But they are a good trade off : LSTMs are thousands of times faster, with only slightly worse results for sequence labeling or production (ie. text and audio comprehension). CNNs are millions of times faster at image comprehension than a fully connected network, and are only slightly worse at it.
So the thing CNNs are really good at are perfect for small images, but symbols already have that thing done. When CNNs and RNNs are used together (e.g. image segmentation), it's almost as if the CNN is creating symbols and RNN s processing them.
Or, you know, maybe I'm far off. I'm in no way qualified to talk about this (just a fan).
Meanwhile, the hardware in our heads is far more complex and capable than the models we're building. And we're not even sure how that hardware works, or what principles we can abstract out and simplify. A neuron getting electrical and hormonal signals is far more complex than an lstm cell... Does it need to be? Or are there biological requirements built in which we don't understand? And how many evolutionary accidents which never evolved away?
It's also possible to go the other way around. PixelRNN uses a kind of special convolutions to encode and generate images. And CNN's have been used for text translation and speech synthesis, especially because they are faster.
The RNN and CNN contain prior knowledge about the domain. The RNN contains the idea that prior elements influence the following elements in a sequence. CNN's encode the idea that close pixels are related. Both imply some kind of domain knowledge.
Multiple attention heads could potentially learn domain knowledge from data.
Whilst there are "standard" approaches in computer vision ("CNNs applied to <foo>") and sequence processing ("LSTM RNNs applied to <foo>"), there doesn't seem to be any "standard" for variable-size, recursively-structured data. Sure there's recursive ANNs, backpropagation-through-structure, etc. but they all seem like one-off inventions, rather than accepted problem-solving tools.
It's one click to get the pdf from there. But you also get a plain webpage with abstract, citation details, and so on, which you can't get back to from the PDF.
In general it's good to knock the ".pdf" off the end of all papers.nips.cc links. Similarly turn /pdf/ links on arXiv into /abs/ links, and replace "pdf?" in openreview.net links with "forum?".
The dominant sequence transduction models are based on complex recurrent or convolutional neural networks that include an encoder and a decoder. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train. Our model achieves 28.4 BLEU on the WMT 2014 English- to-German translation task, improving over the existing best results, including ensembles, by over 2 BLEU. On the WMT 2014 English-to-French translation task, our model establishes a new single-model state-of-the-art BLEU score of 41.0 after training for 3.5 days on eight GPUs, a small fraction of the training costs of the best models from the literature.