For more modern material, there are a few new good books on Transformers. Transformers are interesting because they were designed for efficiency: layers the same size, encoding both data and time sequencing information in each sample (so recurrent NNs aren’t required), etc.
For a higher-level conceptual view of how Transformers work, you can check out the now-classic "Illustrated Transformer" series [3] and this programmer-oriented explanation (with code in Rust) from someone at Anthropic [4].
[1] https://www.oreilly.com/library/view/natural-language-proces...
[2] https://github.com/karpathy/minGPT
[3] https://jalammar.github.io/illustrated-transformer/
[4] https://blog.nelhage.com/post/transformers-for-software-engi...
Published in 1996, but the content is all foundational.
The "deep" part of DNNs has basically thrown mathematicians and statisticians into an infinite loop that they can't quite compute yet. It's a brand new world and we need them to participate.