Generalized Language Models
lilianweng.github.io
lilianweng.github.io
https://ai.googleblog.com/2019/01/transformer-xl-unleashing-...
And for deep background, these recent sets of notes from Stanford's Deep Learning NLP class:
http://web.stanford.edu/class/cs224n/
https://www.youtube.com/playlist?list=PL3FW7Lu3i5Jsnh1rnUwq_...
The original paper [1] wrote:
"All of the parameters of BERT and W are fine-tuned jointly to maximize the log-probability of the correct label."