- Encoder based models have much faster inference (are auto-regressive) and are smaller. They are great for applications where speed and efficiency are key. - Most embedding models are BERT-based (see MTEB leaderboard). So widely used for retrieval. - They are also used to filter data for pre-training decoder models. The Llama 3 authors used a quality classifier (DistilRoberta) to generate quality scores for documents. Something similar is done for FineWeb Edu
So I’m trying BERT models out :)
Also for analyzing Trump's tweets (from 2016): https://mathematicaforprediction.wordpress.com/2016/11/21/te...
See the Mathematica code in this Markdown file: https://github.com/antononcube/MathematicaVsR/blob/master/Pr...
- Generative model outputs are not always desirable, and often even undesirable
- BERT models are smaller and can run with lower latency and serve larger batches with lower vram requirements
- BERT models have bidirectional attention, which can improve performance in many applications
LLMs are “cheap” in the sense that they work well generically, without requiring fine tuning. Where they overlap with BERT models is mostly that they may work better in low training data environments due to better generalization capabilities.
But mostly companies like them because they don’t “require” ML engineers or data scientists on staff. For the lack of care given to evaluation that I see around LLM apps, I suspect that’s going to prove to be a faulty premise.
The most recent version of Wolfram Language (aka Mathematica) uses by default BERT models for embedding.
(Say, for this function: https://reference.wolfram.com/language/ref/CreateSemanticSea... .)