Indigenous engineers are using AI to preserve their culture
nbcnews.com
nbcnews.com
So far I have not had that much luck getting the models to learn the Kiksht grammar and morphology via in-context learning, I think the model will have to be trained on the corpus to actually work for it. I think this mostly makes sense, since they have functionally nothing in common with western languages.
To illustrate the point a bit: the bulk of training data is still English, and in English, the semantics of a sentence are mainly derived from the specific order in which the words appear, mostly because it lost its cases some centuries ago. Its morphology is mainly "derivational" and mainly suffixal, meaning that words can be arbitrarily complicated by adding suffixes to them. So baked into English is word order that sometimes we insert words into sentences simply to make the word order sensible. e.g., when we say "it's raining outside", the "it's" refers to nothing at all—it is there entirely because the word order of English demands that it exists.
Kiksht in contrast is completely different. Its semantics are nearly entirely derived from triple-prefixal structure of (in particular) verbs. Word ordering almost does not matter. There are, like, 12 tenses, and some of them require both a prefix and a reflective suffix. Verbs are often 1 or 2 characters, and with the prefix structure, a single verb can often be a complete sentence. And so on.
I will continue working on this because I think it will eventually be of help. But right now the deep learning that has been most helpful to me has been to do things like computational typology. For example, discovering the "vowel inventory" of a language is shockingly hard. Languages have somewhat consistent consonants, but discovering all the varieties of `a` that one can say in a language is very hard, and deep learning is strangely good at it.
> The last fully fluent speaker of Kiksht, Gladys Thompson, died in July 2012
Which is from https://web.archive.org/web/20191010153203/http://www.opb.or...
Maybe there are some sources talking about fully fluent people still being alive? As currently the article gives the impression they were the last person to "fully" speak it.
That's a fairly common language feature; such languages are generally called "agglutinating".
Prominent examples of agglutinating languages are the Eskimo languages, Turkic languages, and Finnish.
There should be no shortage of resources available if you want to learn Turkish or Finnish.
I am also working on low-resource languages (in Central America, but not my heritage). I see on Wikipedia [0] it seems it's a case of revival. Are you collecting resources/data or using existing? (I see some links on Wikipedia).
re: revival, the Wikipedia article is a little misleading, Gladys was the last person whose first language was Kiksht, not the last speaker. And, in any event, languages are constantly changing. If we had been left alone in 1804 it would be different now than it was then. We will mold the language to our current context just like any other people.
HN is still one of the few places on the internet to get such esoteric, expert and intellectually stimulating content. It's like an Island where the spirit of 'the old internet' still lives on.
While I am not indigenous, I hope to help alleviate this problem. I'd love to hear about your research!
Lots of languages, even Indo-European languages, have very different word order from English or a much less significant word order.
They say language shapes thought. Having an LLM speak such a different language natively seems like it would uncover so much about how these models work, as a side effect of helping preserve Kiksht. What a cool place to be!
https://github.com/australia/mobtranslate.com/
In it's current iteration the homepage is just running dictionaries through OpenAI. (my tribes dictionary fits in a 100k context window)
My old ambitions can be found somewhat here -> https://github.com/australia/mobtranslate-server
That being said, the OpenAI models do a fantastic job at translating sentences so I've put my own model research further to the back. (will try to find some examples)
I can't speak of the true preservation, not many native speakers left, but in my mind that's not even all that important from a personal/cultural perspective.
If the youth who are interested in learning more about their language have a nice interface with 70-80% accurate results and they enjoy doing/learning it then that is a win to me. (and kind of how language evolves anyway) (the noun replacement seems to work great, but grammar is obviously wishy-washy)
(At this point, I just rushed to get my tribes dictionary crawlable so hopefully it will be in a few models next training phases)
Say what you will about llms being overhyped but this is the original core use case.
To me the beauty of these things is in their liveliness, in the aspiration to flourish and grow, not merely to conserve a little longer, to spend one more night with terminal cancer before the inevitable.
The content of modern culture is too much for dying or ancient languages, and what you actually get is English/modern thoughtspace expressed in the lexicon of an until-now separate culture. This flood of spam destroys what was unique and interesting about the culture, and "skin-suits" it.
Anyway, in resurrected ancient languages, this modern samey-ness is a problem. Go read Latin Wikipedia, for example. It's much more like reading modern English authors (though with Latin words and grammar) than it is like reading classical authors.
--
I add a natural language translation (from ChatGPT) of your statement for anyone who doesn't share your lexica:
> "Maybe, unless some version of the Sapir-Whorf hypothesis is correct, and language—along with its grammar, sounds, and other aspects—creates an inescapable worldview that can't be dominated."
By "inalienable unwelt," I mean it in the sense of "inalienable rights" that cannot be removed by external circumstance.
By "inaccessible to hegemony," I mean an umwelt which cannot be perceived by speakers of the dominant language.
The next few decades are going to be really, really weird.
Hallucinations rarely make up invalid grammar or invent non-existent words, what we're concerned about is facts, which isn't relevant at all when the goal is language preservation.
I've seen LLMs hallucinate nonexistent things in programming languages. It's hard to believe it won't do the same to human ones.
If a language model were asked what the word for "floppy disk" was in an extinct language and it invented a decent circumlocution, I don't think that would be a bad thing. People who are just engaging with the model as a way of connecting with their cultural heritage won't mind much if there is some creative application, and scholars are going to be aware of the limitations.
Again, the misapplication of language models as databases is why hallucinations are a problem. This use case isn't treating the model as a database, it's treating the model as a model, so the consequences for hallucination are much smaller to the point of irrelevance.
An LLM should do fine with that since it's usually the foreign word spelled in a way that makes sense in that language. I'm more curious about the inverse though. It's sometimes quite difficult to explain the meaning of a word in a language that does not have an equivalent, be it because it has a ton of different meanings or because it's some very specific action/object.
I suspect it'll be hard to find more material in some obscure, dying language than there is of either of those in the common training sets.
Sometimes they do a decent translation, too often they trip up on grammar, vocabulary or just assuming that a string of bytes means the same thing always. I find they work best in extremely formal settings, like documents produced by governments and lawyers.
The only way this project is going to make sense will be to train it fresh on text in the language to be preserved, in order to avoid accidentally corrupting your model with English. If it's trained fresh on only target language content, I'm not sure how we can possibly generalize from the whole-internet models that you're familiar with.
To me it doesn't make sense. It seems like an awful way to store information about a fringe language, but I'm certainly not an expert.
and getting translation in some direction or other back.
This seem to make a lot of English speakers upset, that LLM outputs appear translated from perspectives of primarily non-English speakers. But hey, it's n>=2 even at HN now.Since LLM:s work by probabilistically stringing together sequences of words (or tokens or whatever) I don't expect them to become fully fluent ever, unless natural language degenerates and loses a lot of flexibility and flourish and analogy and so on. Then we might hit some level of expressiveness that they can actually simulate fully.
The current hausse is different but also very similar to the previous age of symbolic AI. Back then they expected computers being able to automate warfare and language and whatnot, but the prime use case turned out to be credit checks.
It's not exactly nonexistant outside of them, but they make it worse than it is
I find they struggle a lot with things like long sentences and advanced language constructs regardless of the natural language they try to simulate. When it doesn't matter it's useful anyway, I can get a rough idea about the contents of documents in languages I'm not fluent in or make the bulk of a data set queryable in another language, but it's like a janky hack, not something I'd put in front of people paying my invoices.
Maybe there's a trick I ought to learn, I don't know.
Machine learning techniques are really, really good at finding statistical patterns in data they're trained on. What they're not good at is making inferences on facts they haven't been specifically trained to accommodate.
No doubt, it's excellent for archiving, but that's not the same as "preserving" culture. If it's not alive and kicking it's not a culture IMO. You see this happen even with texts : once things start being written down, the actual knowledge tends to get lost (see India for example).
This "AI to help low-resource languages" thing is a big deal in India too, but it just feels like another "jumla" for academics/techbros to make money. I mean, India has brutal/vicious policies that are out to destroy any and every language that's not English (since it's automatically a threat to central rule from Delhi), but pretty much no intellectual, either in India or the US, actually cares about the mass-wiping out of Indian languages by English... Not even the ones who go "ree bad British man destroyed India" on twitter all day.
Learning a new language is a good use case for LLMs, not just indigenous languages, but any language.
As for your comment "ree bad British man destroyed India" this sound more like politics than anything of substance.
(There may be a good answer to that question: perhaps, for example, the corpus can't be preserved for data protection reasons but the LLM trained on it can be preserved? For various reasons that doesn't seem very plausible, however.)
I think of cultures and languages as tools. When they don't serve a purpose anymore, they should be replaced by something more functional.
Cultures and languages die out because they're slowly (or revolutionarily quickly!) replaced by another. It's not like there are people out there speaking no language because their mother tongue has died out.
And who gets to decide that a language has no purpose?
I don't like the way this argument goes.
The parents of the child doing the language learning or the adult doing it.
If and when that's a voluntary and organic process, sure. The problem is that replacement more often that not comes about through ethnic cleansing and violence. These indigenous languages were perfectly functional to the people who spoke them at the time but they were replaced because they didn't serve the purpose of colonizers.
And "lost" languages do get reclaimed from time to time. Hebrew and Irish being two examples.
Hebrew was brought back from the dead for the Jewish refugees populating Israel to have a common language. This solved a genuine practical problem.
Teaching Irish to kids who all speak English does not help anyone communicate better. It seems like a nationalist pride project, and those are not my favorite.
Okay colonizer.
That’s surprising and seems different than what I’ve seen for other languages in other parts of the world (even if it’s a relatively new phenomenon).
Curious, what do you mean by this?
> pretty much no intellectual, either in India or the US, actually cares about the mass-wiping out of Indian languages by English
Well I've never heard of this so lack of awareness would be an obvious cause if it's an issue, are there any orgs raising awareness of it? Also seems surprising to me, Bollywood movies are immensely popular and are all in Hindi. Is there a danger of English overtaking Indian society to the extent where Bollywood movies would mostly be made in English?
So, although all distinct ethnicities may originate from specific places and times, indigeneity as a political and social identity is meaningful only in the context of colonial domination and resistance.
Not every engineer is engaged in an ongoing struggle for sovereignty against colonizing powers.
I think a better question might be what, specifically, does AI bring to the table that justifies its inherent inaccuracy? Why not just use audio recordings, etc?
Used judiciously as a model, there's no problem. They only become a problem when people try to treat them as a database.