BERTs Are Generative In-Context Learners
arxiv.org
arxiv.org
We found Google's T5 models which were released in 2019, pre-GPT-3, were "secretly" capable of in-context learning with a simple inference technique.
Given they use a bidirectional MLM (Masked Language Modeling) objective, it wasn't obvious how to do it, but MLM objectives are known to produce better language representations than causal (next token prediction) objectives. We were able to outperform much larger sized GPT-3 models or get very close to their performance with far smaller T5 models.
While the technique requires multiple samples to coax generations from this particular model, other LLM training schemes have incorporated both unidirectional and bidirectional objectives in their training now. However, this exploration hasn't been fully resolved as most models are still trained only on the causal objective by standard practice. There's still a a lot of exploration that can be done on pre-training objectives.
I think you have to consider that in 2020/2021 many PhDs and Professors attempted to shift grant funded research with BERT and T5 to explore how they could compete with GPT-3 or to express other properties of it that supposedley outdid GPT-3. Very few (besides sentence transformers) succeeded. It's not like this is an unexplored niche. A lot of people in denial were trying to keep on with BERT research for a while despite the fact their work was essentially made obsolete by GPT-3.
(and notably Table 1 and Figure 4 are cherrypicking the smallest size with the largest gaps in task difference, and a size we know decoding is not performative at - 1.3B param mark - the characteristics and conclusions the authors come to (wow, BERT is trained on less data but does better!) obviously can't be made at larger sizes because the actual GPT models become much larger)
I'm having trouble understanding whether this paper is saying anything new. The original BERT paper already compared it favourably to causal models including GPT. Was there any doubt that BERT-style models could be in-context learners?
From what I gather as a non-expert, the problem with BERT is scaling/training efficiency: GPT gets C-1 training examples out of a training input of length C, but BERT only gets 0.15*C examples. Indeed, the author points out that DeBERTa required 3x more compute than GPT-3 to achieve the level of performance reported, which makes sense.
Also for analyzing Trump's tweets (from 2016): https://mathematicaforprediction.wordpress.com/2016/11/21/te...
See the Mathematica code in this Markdown file: https://github.com/antononcube/MathematicaVsR/blob/master/Pr...
So I’m trying BERT models out :)
- Encoder based models have much faster inference (are auto-regressive) and are smaller. They are great for applications where speed and efficiency are key. - Most embedding models are BERT-based (see MTEB leaderboard). So widely used for retrieval. - They are also used to filter data for pre-training decoder models. The Llama 3 authors used a quality classifier (DistilRoberta) to generate quality scores for documents. Something similar is done for FineWeb Edu
- Generative model outputs are not always desirable, and often even undesirable
- BERT models are smaller and can run with lower latency and serve larger batches with lower vram requirements
- BERT models have bidirectional attention, which can improve performance in many applications
LLMs are “cheap” in the sense that they work well generically, without requiring fine tuning. Where they overlap with BERT models is mostly that they may work better in low training data environments due to better generalization capabilities.
But mostly companies like them because they don’t “require” ML engineers or data scientists on staff. For the lack of care given to evaluation that I see around LLM apps, I suspect that’s going to prove to be a faulty premise.
The most recent version of Wolfram Language (aka Mathematica) uses by default BERT models for embedding.
(Say, for this function: https://reference.wolfram.com/language/ref/CreateSemanticSea... .)
But I believe we’re confident that all major models are causal transformer models right now.
No reason to believe otherwise. If one of them was doing something different, they’d let us know in order to stand out.
“Transformer” is the name of the algorithm behind popular LLMs.
GPT is the name that openai gave to their models early on.
The Otis company invented the term escalator, and even had a trademark on it for a while, but does it mean that you'd only call one an escalator if it was made by them?
Why try arguing this?
They make impressive demos, but I can't recall any of their released models being at the top of any leaderboard.
EDIT: Sorry, looking into it a bit more now, they still seem to be at the top in term of the context window, so they got that going for them.