Electra: Pre-Training Text Encoders as Discriminators Rather Than Generators (2020)
arxiv.org
arxiv.org
Unfortunately BERT models are dead. Even the cross between BERT and GPT - the T5 architecture (encode-decoder) is rarely used.
The issue with BERT is that you need to modify the network to adapt it to any task by creating a prediction head, while decoder models (GPT style) do every task with tokens and never need to modify the network. Their advantage is that they have a single format for everything. BERT's advantage is the bidirectional attention, but apparently large size decoders don't have an issue with unidirectionality.
If you're running 100k QPS through the model with a budget of 0.1 cents per query, you aren't going to be using a GPT model for classification.
You can have a bidirectional model directly fill in the middle...
Or you could just frame that as a causal task by giving the decoder llm a command to fill in the blanks, and the entire document with the sections to fill replaced by a special token/identifier all as input, and the model is trained to output the middle sections along with their identifier.
There we go, now we have a causal decoder transformer that can perform a traditionally bidirectional task.
The gains in training efficiency and compute cost versus widely used text-encoding models like RoBERTa and XLNet are significant.
Thank you for sharing this on HN!