Also, if you're going to do a 30k foot view of a technical topic, you might want to tell people what GPT3 is somewhere in there.
Also, if you're going to do a 30k foot view of a technical topic, you might want to tell people what GPT3 is somewhere in there.
I feel the need to defend the author though, it's hard to make research accessible while still distilling valuable insight. I think his post on transformer networks [1] did a good job for example, and you'll appreciate the lack of animations.
In addition to your link, I've found a really good Transformer explanation here (backed by a Github repo w/ lively Issues talk): http://www.peterbloem.nl/blog/transformers
Additionally, there's a paper on visualizing self-attention: https://arxiv.org/pdf/1904.02679.pdf
This article and the animations definitely helped me a lot in understanding this. I learned quite a few things, so thanks a lot to the author!
But these animations/diagrams are so high level that they could be used for Explaining all sorts of NLP models from the past 5 years.