Pg 26, quote: "Anything humans can do in 0.1 sec, the right big 10-layer network can do too". That is a very bold claim. It encompasses the entire fields of image and voice recognition as well as knowledge encoding. It's slowly becoming clear that this is likely to be true.
Pg 39, 40: Google's ImageNet-winning system in 2011 had 7 layers and an error rate of around 16%. The 2014 system had 24 layers and an error rate of 6.66%. Note that trained humans have an error rate of around 5%[1].
Page 50-57 talk about the miracle that is Word2Vec, and what is possible with that.
Page 60-70 talks about paragraph embedding. I haven't seen this published before.
Page 70-73 extends word/paragraph embedding for translation. I've seen a slide deck showing this works before, but I need to read the new paper cited there.
Page 74+ talks about cross-modal embeddings, especially the caption generation stuff. HN has had a few things on that over the past month or so.
[1] http://karpathy.github.io/2014/09/02/what-i-learned-from-com...