Language-Agnostic Bert Sentence Embedding
ai.googleblog.com
ai.googleblog.com
"We adapt multilingual BERT to produce language-agnostic sentence embeddings for 109 languages. %The state-of-the-art for numerous monolingual and multilingual NLP tasks is masked language model (MLM) pretraining followed by task specific fine-tuning. While English sentence embeddings have been obtained by fine-tuning a pretrained BERT model, such models have not been applied to multilingual sentence embeddings. Our model combines masked language model (MLM) and translation language model (TLM) pretraining with a translation ranking task using bi-directional dual encoders. The resulting multilingual sentence embeddings improve average bi-text retrieval accuracy over 112 languages to 83.7%, well above the 65.5% achieved by the prior state-of-the-art on Tatoeba. Our sentence embeddings also establish new state-of-the-art results on BUCC and UN bi-text retrieval."
(Found via https://tfhub.dev/google/LaBSE/1)
Releasing model weights has been common for a long time, it was OpenAI that regressed and refused to release its GPT models until it released GPT-3 hidden behind a semi-public API. The google BERT models are released pre-trained in full (as mentioned in the article), you can easily play around with them on your own.
BERT isn't a forward LM like GPT, so it's a little less easy to play with for the uninitiated - although there are papers showing text generation with BERT.
At least that's my understanding. Feel free to correct me if I'm wrong.
GPT-2 wasn't really feasible for most actors to train from scratch, otherwise it would have been released by a third party. GPT-3 is technically feasible for a single actor to do inference with, although not really really.
Still in the realm of doing inference in homelab territory (barely).
Where's this heuristic from? Seems handy if true.
Edit: It probably is, considering you only need weights to be fp16, while intermediate layers can be fp32 and reuse these memory.
There is also no reflection on why some languages got better results than others. Again looking at the last figure in the top three best performing languages with no data two are Sinitic languages (Cantonese and Wu) which likely happens since their closeness to standard Chinese for which there is huge ressources available for training. On the reverse Breton for which there is probably very few apparented language data in the initial models and training set besides Irish get very poor results, which tends to show the model actually don’t transfer very well if at all...
https://1.bp.blogspot.com/-eqH1ZsnTP2o/XzwlDjy3X5I/AAAAAAAAG...
It is very concerning how few thought is usually put into linguistic or language characteristics when dealing with these topics. I also rarely see cultural considerations etc. Basically everything is considered as "machine learning will hopefully get this right if having enough data" which is unfortunate (ML is a great tool but the conferences are about language processing).
Another big issue I noticed is that a majority of research only targets or evaluates English texts. In many cases the language is not even specified (although it is clear they use English from figures or examples). I even heard people complaining that work on non-English data is treated as too minor by many reviewers so stuff like that often just gets rejected.
I think this is a really weird development for a field which centers around natural languages.
It's also a bit simplified to consider it a bifurcation between "traditional" linguists and AI experts entirely ignorant of the discipline. Long before the current wave of AI started, Google liked to hire linguists and computational scientists. These teams probably do have plenty of subject matter experts, but for now they are reaping the low-hanging fruits of the suddenly-improved generic methods. As the marginal improvements are inevitably diminished, subject matter will become more salient again.
I'm a computational biologist by training, and have great appreciation for the often beautiful algorithms, many created in the 70s or 80s and allowing then-spectacular feats of tackling large datasets. Unfortunately, it isn't always obvious how to transfer that knowledge to the new way of doing things.
I'd argue, that improving the ML models is really the job of ML researchers and should be mainly targeting ML conferences like AAAI (Adv. of AI). In other conferences (directly targeting NLP, CV, Comp. Biology, etc.) it should be the main job to combine those models with the domain-specific characteristics (e.g., language information for NLP) or "traditional" methods to make it an interesting discussion.
I was recently doing reviewing for a multimedia conference and quite a lot of the papers I reviewed were basically pure ML papers. A colleague had the same experience.
I agree that there should be some discussion on why transfer to certain languages is easier than to others (and maybe that's actually in the paper, I've only read the blog post so far). Language relatedness is one obvious explanation, but another possibility is contamination of the training data. It's rare to find Cantonese on the internet marked with lang="yue", much more often it's simply labeled as Chinese, i.e. lang="zh". That makes it hard to collect "pure" Standard Mandarin data that doesn't include other varieties.
There is a reason, which is akin to why balanced corpus are used for capturing various aspects of a language: language features are unevenly distributed between languages but quite consistent within a family.
Let's take the example of word order [1]. The SVO order is very common on the raw number of languages, but actually represented by a small number of families (4 times less families than SOV for about the same number of languages). Which mean if the so-called "agnostic" model is trained mainly on languages of the same family (spoiler: it is), the raw number of languages used can be high, yet the difference between them is minor so the task is made easier.
In ML term the model is kind of overfitting the feature from the language that are the most represented, yet the result can pass as good when tested on languages similar to those used for learning. This is clearly what could be happening here, from the previous example I gave. Also Catalan: 75% accuracy for a language that is linguitically very close to French and Spanish (themselves close together) which are the top 5th and 9th language in number of data in the training set. The first two best non constructed languages are the closest to the 4th language in the training set...
Besides, I don't think anyone really on lang tag as it often missing or simply wrong. Besides, taking Cantonese from Standard Chinese is a very easy task as their is a few words that appears very frequently in Cantonese and never in SM (係, 嘅, 冇 for instance).
[1] https://en.wikipedia.org/wiki/Word_order#Distribution_of_wor...