Attention and Memory in Deep Learning and NLP
wildml.com
wildml.com
There is a lot of neuroscientific work on attention, really a lot! Overt and covert attention. Microsaccades, very small eye movements, with already a bunch of possible functional roles. Almost everything we know about the brains of little kids is by studying where they look at and where they pay attention to.
Structure-wise attention models can be quite simple. The structure that is often seen is a WTA (winner-take-all) network with subsequent serial inhibition. The first winner is inhibited, so the next winner can come on stage. This is the same system as Baars has in his global workspace theory [3]. It is also the same method as in mundane RANSAC models [4]. That's a workhorse of computer vision in which a consensus/voting model can be used to have data points voting for higher-level structures. When one structure is detected, votes for it are removed, and the next most salient structure can be voted for.
[2] http://cns-alumni.bu.edu/~yazdan/pdf/Itti_etal98pami.pdf
I had the same feeling about boundary box recommendations/guesses that were used to speed up object recognition with Deep Learning fairly recently. Just as with a sliding box approach it is intuitive and works, but it also seems quite inelegant and like a better approach should be possible. Visual attention seems like it should work much better in the long term, so it is exciting the field has come to a point where it has been developed.