Just training on existing sounds might give us a whale sound generator. But it would be the equivalent of Prisencolinensinainciusol: https://youtu.be/-VsmF9m_Nt8
Just training on existing sounds might give us a whale sound generator. But it would be the equivalent of Prisencolinensinainciusol: https://youtu.be/-VsmF9m_Nt8
Not Necessary.
Searching for Needles in a Haystack: On the Role of Incidental Bilingualism in PaLM's Translation Capability (https://arxiv.org/abs/2305.10266)
ablation studies performed on the effects of removing tiers of multilingual corpora (parallel corpora, bilingual (not necessarily parallel), and english only)
Removing Parallel corpora or even just bilingual data in general reduces quality but does not make the model unable to translate. Crucially, the gap or reduction in quality also seems to diminish with scale.
The paper also shows models may not need any multilingual corpora at all. The 8b Model can still translate latin scripts with only english training data, adding evidence for larger models being better equipped to leverage either sparse signals (i.e., language-identification failures during ablation) and/or weak signals (i.e., language similarities from shared scripts).
> LLMs do not learn to translate through explicit examples of parallel text.
Of course they do, lots of that is in the training data.There are several completely unsupervised language translation systems that predate LLMs, but the performance is middling.
For LLMs the translation behavior is largely an emergent property that is not completely understood; if we can tokenize whale language in a useful way, it is entirely possible that the LLM can derive a weak or approximate translation of some of the language structure.
Maybe a better example is in the multimodal context. Drawing images from words is a type of “translation”. But for this task we need captions, i.e. a parallel corpus.
However, I believe you are overstating the results of the paper. If anything, the paper demonstrated massive reductions in zero shot translation capabilities after removing translation. And for languages with no cognates with English, BLEU scores are pretty abysmal. So the paper suggests that parallel corpuses still very important.
For whale languages, you won’t have any cognates with English, and we don’t even know if they have a grammar that is remotely human.
Like I said, the crucial thing is that it looks to be a diminishing gap. I'm not saying parallel corpora is doing nothing or is unimportant but that's a clue that monolingual competency of the languages in question is the biggest factor here. The extreme positive transfer language models exhibit in terms of multilingual competency is another clue. https://arxiv.org/abs/2108.13349
Parallel corpora may be a crutch with fading relevance.
Please expand -- as I understand it the LLMs use cross entropy on next token prediction as their loss function and I fail to see how this gives any feedback about translation, even given parallel texts. Predicting the next token in a French language text and predicting the next token in an English language text are not obviously more interesting than dealing with non-parallel texts in both languages.
It's totally possible that these parallel texts are critical in a sense to the translation capabilities but this is not obvious.
It would be incredibly difficult to establish empirically because reducing training data reduces the effectiveness at all tasks. Translation, like almost all of the capabilities of LLMs, is an emergent behavior that we don't fully understand yet.
To add, large enough here is also a potentially shifting target. LLMs exhibit extreme positive transfer of language competency. So 5B tokens of Korean will take Korean competency much further if trained alongside 50B tokens of English than if alone.
What would a whale call a Quarter Pounder with Cheese?