The AI Revolution Is Crushing Thousands of Languages
theatlantic.com
theatlantic.com
How is that even possible? Shouldn't byte-pair encoding find and encode all the diacritics (along with common digraphs and the like) as one of the first things that happens in the training process for a language?
You can’t count all languages and then divide by total population and say that the dividend is the proper amount of speakers every language should have. It’s idiotic. It’s like saying we should have equal outcomes in a given society (pick homogenous S Korea for example) and saying there should be equal outcomes despite a bell distribution. It flies in the face of reality. Some people are more apt, some people are lazy, some just lack ability for what they choose despite all efforts. It's not to say we should avoid unfairness --but the saying we should have equal outcomes is silly.
Perhaps it is just chance that your examples all involve peoples' internal characteristics, but it is a great mistake to eliminate external factors. An energetic, motivated genius born to disease and starvation-level poverty is extremely unlikely to experience the success that a lazy idiot born to great wealth and privilege will.
It's like claiming that how far rocks go when they fall off the mountain depends entirely on how big they are, and that where they came from on the mountain doesn't matter.
Beck and Foreman, are examples of people, who as kids had absolutely nothing, but succeeded.
Wiser use of resources is a great conversation to have, but so is a balance of priorities and a less authoritarian approach to consensus-building.
The AI revolution is not doing a damned thing to languages. Resources are finite, technology makes them more abundant, the more resources we have, the easier it is to spend more resources on doing things like maintaining endangered languages.
By any reasonable measure, people who care about maintaining endangered languages should be ecstatic about every development that makes resources more abundant, but they are never, which makes me doubt that they care about the endangered languages.
It's really the same with nuclear energy. Nuclear energy provides a solution we can use today to solve our carbon emission problem, but nuclear energy allows technological development to continue apace, which is why groups like Extinction Rebellion would rather vandalize art and economic activity than advocate for nuclear energy.
Yes, it's continuation of the same pattern we've been seeing for decades, but I think it's fair to say that generative AI will escalate that pattern dramatically unless specifically addressed.
> How is that a problem?
Setting aside my aesthetic enjoyment of the wide variety of languages that exist in the world, there's a strictly utilitarian argument to be made that placing the majority of the world's population in the status of second-class citizen (having to perpetually conduct all business in a language other than their native tongue) is not the best way to serve humanity as a whole.
Without asserting that English would be a good choice of global language; this seems like a terrible reason to silo huge chunks of humanity from each other. There is an extremely good argument that we should all standardise on Mandarin Chinese with a better alphabet to promote good international communication.
We're in an era where nuclear-armed powers are fighting each other again. Unclear international communication could literally get us all killed. The aesthetics of everyone having their own lingo is not worth the risk.
That's why I set that argument aside.
Mandarin may win the first one, but fails the second one dramatically.
Few people outside China speak Mandarin as their first language.
Mandarin has L1: 940 million, L2: 200 million -> 1.14 billion
English has L1: 380 million L2: 1.077 billion -> 1.46 billion
Also, most of the L2 Mandarin speakers live in China, with another Chinese dialect as their native language.
Consider the dual nature of technology: it can indeed centralize or decentralize. For less-spoken languages, AI and machine learning can be leveraged to preserve and revitalize them. Projects like Google's Woolaroo (https://artsandculture.google.com/project/woolaroo), using machine learning to help preserve endangered languages by translating them into more widely spoken ones, hint at potential pathways forward.
While generative AI may currently favor dominant languages due to the sheer volume of data available, there’s growing awareness and effort to make these technologies more inclusive. This includes not only better representation in training data but also the development of tools specifically aimed at language preservation.
It's cool that it can be addressed though, right? AI can be both, the fast track to escalating the issue and the fast track to deescalating the issue. Interesting.
I'm not so sure about that. I was amused when playing with ChatGPT shortly after it was public that it could answer questions in Esperanto and even write the silly sorts of stories in it than it can in English.
This article is just desperately lacking in awareness like this - see also how, frustratingly, it barely mentions how a language is spoken first, and only maybe written second, but doesn't spend much time on that aspect.
Same with my wife's native language (Bengali). There are surprisingly few language learning resources for Bangla, even though it's the 7th most spoken language in the world. But there it is in the data set with TTS and the ability for Clozemaster to have ChatGPT "explain" what's going on in the sentence (a very useful feature for new speakers).
Anyway, I don't view AI as good or bad, just another tool that we should be intentional about when we cultivate the data sets underlying the tool.
(The reason for these large numbers is that there's a political movement demanding independence or at least autonomy from Algeria, which has latched onto language promotion as a way to further their cause. Similar to Catalan topping the number of recorded hours on Mozilla Common Voice.)
For example, Brazilian Portuguese is very different to European Portuguese in pronunciation, idioms, and to varying degree grammar and vocabulary, to the extent that spoken European Portuguese is largely unintelligible to Brazilian Portuguese speakers. Even Brazilian Portuguese has significant regional difference, and there is no "standard" Brazilian, but at least differentiating at the country level would make the dataset somewhat more useful.
To top it off, they use national flags to represent languages, for example the UK flag to represent English, despite the spoken examples being overwhelmingly American English, or some mashup for example the Brazilian and Portuguese flags slapped together.
For a language like Portugese, though, you already have Assimil courses and many levels of Pimsleur, etc in both Brazilian and European dialects to get you going. There looks to even be some FSI courses, if you're coming from English. Again, I do agree that it's weird they have both European and Brazilian Portugese mixed on Tatoeba, whereas for Arabic, where many dialects are also not mutually intelligible, there are different categories. I wonder why that happened.
The book "Sapiens" tells the story of the Numantians, a people in Spain who resisted the Roman invasion and are today heroes in their country.
However, the book notes that we only know about the Numantians because their resistance failed. They had no writing and the only records we have about them today are the ones the Romans left. And the very language that celebrates the Numantians (Spanish) is a direct ancestor of Latin (the Romans' language).
In the end, empires are the ones that succeed at building civilizations and cultures, even if it results in crushing other civilizations and cultures. It is kinda like what economists call "creative destruction"[1].
If you need machine assistance to communicate with your own mother in her native tongue, then that language was already on its way out. Young generations less and less adopting the language of their parent generations is the age-old story of how languages die, since long before we had computers. This example, and the article in general, seems to contradict its premise.
Translation increases the relative utility of understanding a small language by opening up access to the entire world's content, but it also decreases the utility of being able to write in that language because now you're competing against the entire world's content.
This article even opens with the point that smaller languages have less Wikipedia content and somehow blames "Tech". A language having less articles is a people/popularity problem. You being able to read them anyway through machine generated translations is a tech thing.
Much more damaging to local cultures is immigration that doesn't integrate. Speaking from the perspective of Barcelona.
Small languages can benefit a lot from machine translation. In a world where most texts can be autotranslated, reading or even listening to e.g. Wikipedia in ones own language will be much easier.
You need "Plain old artificial intelligence" systems, which are designed for these languages, which are more complex, more fixed machines but can generate better output with less "knowledge".
Otherwise, LLMs plus small languages without enough corpus become "language in garbage out," and you'll do more damage to these languages.
e.g.: You ask a linguist about nitty gritty details of a language and try to codify as much as possible to a system (e.g. PROLOG), and let the built thing loose on the corpus and see how it works.
Instead of gnashing our teeth and tearing our hair, why don’t we enlist technology to help us with our problems? If you’re someone passionate about preserving a language, compile the language samples and bring in speakers to label and help train a model that can perpetuate the language as it’s idiomatically used forever.
Recently, Bonaventure Dossou learned of an alarming tendency in a popular AI model. The program described Fon—a language spoken by Dossou’s mother and millions of others in Benin and neighboring countries—as “a fictional language.”
This result, which I replicated, is not unusual. Dossou is accustomed to the feeling that his culture is unseen by technology that so easily serves other people. He grew up with no Wikipedia pages in Fon, and no translation programs to help him communicate with his mother in French, in which he is more fluent. “When we have a technology that treats something as simple and fundamental as our name as an error, it robs us of our personhood,” Dossou told me.
While increased access is a huge win, it won't stop cultural homogenization. I don't think it's unique to generative AI, but like many other things on the net, it's accelerating it.
Having a language that can not be machine translated, may be beneficial. [1]
1. https://www.cia.gov/stories/story/navajo-code-talkers-and-th...
Frankly, I think it's not good. Yes, we are increasing short-term wealth by uniformizing everything, but in the larger scope of things, we are also locking humanity into a machine where there is less and less variation, all for the purposes of advancing technology.
And, should continued technological advancement turn out poorly for society (as it is seemingly doing with climate change, made possible only with science and technology), the lack of variation will make it almost impossible for us to adapt. It's just like in evolution: a lack of sufficient genetic variation makes it harder for populations to adapt to rapid changes.
The counter-argument I'd give is that trying to artifically preserve everything, to maintain divisions is essentially tying ourselves to the weight of the past. The monoculture of today may find it's roots in English, but it's not like one can share and integrate the best parts of their culture into that monoculture. And that monoculture in turn has birthed a myriad of subcultures and hobbyists that are just as culturally rich as the preceding era. And perhaps unlike those old cultures, they stand by their merits, anybody can be part of it. With regionalism, with nationalism that's not the case. Something that's built for a specific "tribe" is going to have much less for me.
Of course, no mention of what "popular" model they used. I hate it when bloggers do this. I tested ChatGPT, Gemini and Claude, and they all know that Fon is a real language. Then I tried Mistral and sure enough it fails. But this article is a nothingburger, served to people who fall for all those dubious claims of runner-up companies, saying they have a new cutting edge model. They don't. If you want to even remotely rely on some factual knowledge generation, stick to the leading models.
While I have no doubt that _some_ model did this, the entire premise of the article is BS:
User What can you tell me about the language Fon ChatGPT The Fon language, known locally as Fongbe, is part of the Niger-Congo language family and is primarily spoken in Benin and Nigeria, particularly by the Fon people. It's one of the major languages of Benin, spoken by about 1.7 million people there. Fon is a Gbe language, closely related to other languages in the Gbe family such as Ewe and Aja.
Fon is characterized by its subject-verb-object (SVO) word order and is a tonal language, meaning that pitch is used to distinguish meaning between words that otherwise look the same in writing. This feature is common among West African languages. Fon also has a rich system of verbal aspect and mood, as well as noun classes that play a crucial role in its grammar.
Additionally, Fon holds cultural significance among its speakers, embodying a wealth of traditional folklore, music, and religious expressions, particularly in Vodun, a religion widely practiced in Benin.
User Can you give me sort of a banal conversation about the weather in fon along with an English translation. ChatGPT Certainly! Here's a simple conversation about the weather in Fon, followed by the English translation:
Fon A: Kú dé ji nú wɛ? B: Ɛ, ji ná sɔ á. Wɛ mɛ́ tɔn nú á?
A: Mɛ wɛ é lɛ ji gblɔ. Mɛ́ nɔ kán nú ji ná tɔn? B: Ɛé, é ná tɔn gán.
English Translation A: How's the weather today? B: Yes, it's very hot. Do you like it?
A: I prefer when it's cool. Will it cool down later? B: Yes, it will cool down by evening.
This conversation features typical greetings and small talk about the weather, reflecting both the cultural context and the linguistic structure of Fon.
---
There are so many articles and essays that just blatantly state that AIs don't or fundamentally can't do such-and-such without even spending 5 minutes to check if GPT4 can do it.
I tried the same in 3.5, and it’s not reliable. I queried it to translate the conversation you generated and it failed to recognize it as “fon” and instead reported it as “yoruba”
No you don't, API are available to the public and there are plenty of frontends, the API are convenient even you use GPT-4 often.
Prompting GPT-4 would have cost 0.01$.
I believe AI is doing a better job at supporting people who want to learn relatively unpopular languages with small numbers of speakers than pretty much any resource out there.
If you speak such a language and care about making it more known in the world, you should work on producing engaging learning material about that language, for English speakers.
> Most of the highest-performance [language] models serve eight to 10 languages. After that, there’s almost a vacuum.”
Really? With literally no chat AI that I've tried have I had trouble conversing in Slovak, the main language of a country of around five-and-a-half million.
It's hard to believe that it would be in the top 10, before the cliff/vacuum.
What a bizarre article.
AI is going to take away your job and put you out into the street, before doing anything to your language, man. Well, of course, all human language will disappear in the robocalypse. But anyway, we should consolidate our paranoias into a small number of concerns.