Using AI to decipher languages that remain a mystery to us
restofworld.org
restofworld.org
But then there's Linear A. Only has 1500 inscriptions or so. And no translation.
So, Harappan/Indus script has about 4K known inscriptions. And descendant languages to work backwards from are still unknown.
Well, I wish them good luck. But I don't know if ML can do any better than humans. It might be able to do what humans can faster, but I doubt it can do what we can't.
That said, I would love to be proven wrong, because I love the idea of the Indus valley civilisation being where "all the wise men in boats" (to paraphrase Pratchett) came from, given the obvious proximity of the Indus Valley to the Persian Gulf and thus Mesopotamia.
Bonus points if someone translates a Harappan inscription and it says "By the way, if some guy called Noah turns up in a boat, claiming to have been spared from a global flood, he's a bit nutty, it was just a heavy monsoon."
Being faster might be enough? Who knows, perhaps some human spending a few centuries in deep contemplation of Linear A might crack it?
However, that doesn't mean that trying to solve it with ML is completely useless. If you can do a century's worth of hypothesis-testing in a fraction of the time to discover that none of them work, at least humans will waste less time trying to solve an unsolvable problem.
If I write down a sentence in a language you don't speak, and make it impossible for you to ever get any information at all about the language it's written in (say, by burning every book about it and killing everyone who knows it), how many centuries of contemplation do you think it would take for your to understand it?
Language decipherment, whether by human or machine, is not done by staring at the words you don't understand. It's done by comparing it to things you do. And here we have nothing.
To reply to your other comment, rabbit and rocks would occur in different contexts, but you have no way of knowing which is which. You can get numbers, probabilities and mutual information scores and vectors and what have you, but if you can't ground them in anything, you have no way of getting to meaning, and therefore no way of translating.
But you have a grounding! You know it's a language used by humans.
It translates very well actually. I've tried with English/French and English/German with very good results.
I think it is quite well understood that deciphering a language without external information is possible? But GPT-3 was a recent example that I think proves it. There are probably tons of literature about this question. I think aliens would be able to understand lots of things about Earth just by listening in.
> To reply to your other comment, rabbit and rocks would occur in different contexts, but you have no way of knowing which is which. You can get numbers, probabilities and mutual information scores and vectors and what have you, but if you can't ground them in anything, you have no way of getting to meaning, and therefore no way of translating.
Well, GPT-3 translates very well, so this argument can not be correct. I don't see why it would be like that? It seems quite intuitive that it would be possible to extract meaning.
> I think it is quite well understood that deciphering a language without external information is possible?
I don't think this is understood at all, and I would have to ask you for just one example of it. I would, in fact, say it is categorically impossible, if by "no external information" you mean that we have only text or sound.
The reason that it is not possible to extract meaning is that meaning is a substance that is not present in the text. It is more than the relations between words, in that it is connected to the world.
The key is, if you have aligned or parallel texts, then you can use structural similarities in the lexicon to bootstrap a mapping from one language to another without referring to substantive meaning (which I will maintain is how humans do translation), even if you don't have a dictionary as such. This is not surprising, but I would expect it to have severe limits.
One of these languages, I've given meaning to the words, the other is random. However I create massive corpuses for each of hundreds of thousands of documents. Now go ahead and machine learn the hell out of both languages.
First question : which one is the 'random' language? Second question : what is the meaning of the non-random language texts? Assume I have not based my 'real' language on any existing example ...
There are principles underlying all human languages. For example, they are used to describe the world or communicate intention and they must be understandable by a human.
Given that the copious data of the non-random language should describe something we understand about the world or social dynamics, we can try to match them with world/domain models and find patterns of plausible interpretations that fit.
The patterns should also satisfy the principle of communication economy and the limits of human cognition, such as the size of working memory. These constraints help narrow down the possibilities.
The human/AI system can test and figure out the most likely patterns first and use them to decipher more patterns.
All above is possible in principle but beyond the current state-of-the-art.
Note that this doesn't mean that relationships between words are arbitrary, they are of course highly systematic, else language wouldn't work. It also doesn't mean that language doesn't have history, the changes are also systematic. But if you take a random surface form (sound or writing), there is absolutely no way to relate it with any certainty to any content without a context that allows you to infer it.
Yes, it's hard to figure out cow just from a big corpus; though you might learn enough about 'cow' to form meaningful sentences and answer questions. That's what GPT-3 is doing after all, and language models are still improving quickly.
And we do have context! Just knowing that a language is used by humans gives lots and lots of context. Human cultures and languages have more in common than you might realize.
For example, have a look at Deixis (https://en.wikipedia.org/wiki/Deixis):
> In linguistics, deixis (/ˈdaɪksɪs/, /ˈdeɪksɪs/)[1] is the use of general words and phrases to refer to a specific time, place, or person in context, e.g., the words tomorrow, there, and they. Words are deictic if their semantic meaning is fixed but their denoted meaning varies depending on time and/or place. Words or phrases that require contextual information to be fully understood—for example, English pronouns—are deictic. Deixis is closely related to anaphora. Although this article deals primarily with deixis in spoken language, the concept is sometimes applied to written language, gestures, and communication media as well. In linguistic anthropology, deixis is treated as a particular subclass of the more general semiotic phenomenon of indexicality, a sign "pointing to" some aspect of its context of occurrence.
> Although this article draws examples primarily from English, deixis is believed to be a feature (to some degree) of all natural languages.
I suspect with a large enough corpus you can most likely figure out the deixis of the given language.
See https://linguistics.stackexchange.com/a/4046 for some more structure shared with all human languages:
> However, all languages have some sort of a clause-type thing allowing them to express predication, attribution, etc. See Dixon's Basic Linguistic Theory: Basic Linguistic Theory Volume 1: Methodology . All languages also must have means of expressing cohesion and coherence (texture) although this is much less studied in cross linguistic perspective. Punctuated sentences are a kind of cohesive device.
(A comment on this answer points out that the reality is probabilistic, of course:)
> While I agree with you, it seems to me that linguists who study languages with a strong written tradition often talk about 'sentences', even when their examples are not from writing. Those of us who work on previously unwritten languages (I think) tend to talk about 'texts' or 'utterances'. I think this is an important issue as it relates to how some linguists have ignored the true messiness that is often found in examples of human language.
Of course, all of this refers to natural human languages. Not totally arbitrary constructs.
As I've said in a sibling comment, GPT-3 is able to construct new utterances in the same language that it learns from. It does not learn its language while learning a model of the world because it does not exist in the world, so it does not learn semantics like we do, which is requirement for making a translation. I am familiar with all these things you talk about, but you are not starting from a sound understanding of semiotics and the difference between content and expression.
Say you learn a lot of syntax from text, you are able to write out a list of paradigms. How do you know which verb form is past and which form is present? How do you know which word means day and which word means night? The information required to make that judgement is not present in the text artifact.
There's lots and lots of context. One interpretation will fit the data much better than another.
Eg in spoken languages the equivalent of 'I' is used much more often than any other pronoun. Similarly simpler sentences are perhaps more likely to refer to the present than the past; eg past tense is much more likely to come with an additional specification of _when_ stuff happened. That information is less kind of a given when we are talking about _now_.
I agree that all these connections are only probabilistic and you need enormous amounts of data.
Btw, GTP-3 can translate between English and French to certain very limited extent; but that's probably mostly because it has seen example translations.
To give you a testable prediction of my theory: if you trained something like GPT-3 on a corpus that includes examples of natural languages A and B, but never any piece of text that contains both languages, yet alone example translations, I would expect that nevertheless, the resulting network would share the same internal representation for the concept of 'cow' in both languages.
This still rests on a lot of assumptions about the language. Pronouns are used much more in some languages than others, some languages have more pronouns than others, etc.. Tense is not at all universal in the world's languages as you probably know. There are linguistic theories that even question the universality of nouns and verbs as categories, and even if they are to an extent universal, they are certainly not fixed in their boundaries across languages, e.g. most things that in English adjectives are expressed as verbs in Classical Arabic. Even if you could do a perfect job of extracting a set of syntactic categories and morphological paradigms from a language, I don't think you would be able to do anything with it. Word 3849 from category C often combines with words 201, 635, and 9913 from category F in conjugations 1-4, but never 5.
And of course, even for nouns that refer to basic, physical things, often some languages will have several words where another will make do with one, or have none. It quickly becomes a question of culture.
Whether you can go from this statistical analysis to anything with absolute meaning, I don't know. I suspect _some_ extremely common concepts (pronouns, verbs like to be or to have) might yield to this sort of direction.
All (bar none) decipherments of unknown writing systems have been done by reference to related languages, especially the existence of bilingual texts. Computers and statistics are definitely able to help us do that, but if all you have is a page of text with no context, then the information you're looking for is gone.
However, I do strongly disagree with the other point in the comment I was responding to, that you could not tell a real language with meaningful words from one with random words. That (entirely made up and not very helpful) problem does seem like it would be amenable to basis statistical analysis.
The reason the first question is relatively easy is because of syntax (sentence order, but also the order of words in phrases). If it's a real language, that order is not random; the English word "the" shows up before "man" much more often than it does before "listens", and so for other words. This is true even in so-called free word order languages, which are really free phrase order languages.
The other thing that could make the first question relatively easy is morphology. If you've decided to make a language that has inflectional morphology (in English, suffixes like -ed and -ing), that also helps determine whether the language is "real", because in a real language not all words will take the same affixes--you can't say 'theseing' (these-ing) in English. In other words, the affixes that you find on a given word are not random.
Of course if you decide not to give your language inflectional morphology, this won't help--but if you do that, then your language's syntax will have to be more rigid: in a language with case marking affixes, like Latin, you don't need word order to tell which noun is subject and which (if any) is object; but you need a more or less fixed word order in English, because (apart from pronouns) we don't have any case marking.
I was only elaborating the more general point that in practice merely being faster can let you do new things.
-- John Tukey
When I'm finished, I wanted to play around with some ML exercises to see if I could derive anything useful in understanding it.
Besides this article does anyone know of any practical/simple attempts at running ML over historical languages?
About dictionary making: there are several programs "out there" that are already set up for making dictionaries, probably better than going from scratch with YAML. SIL's FLEx is very good; their older Shoebox program is also useable, but like the C programming language, you can shoot yourself in the foot. There's another one out of East Africa, but I can't remember its name.
I don't know much about Australian languages, but I think most of them have a lot of morphology. If yours does, then you'll want to do something about inflected forms of words (think English walk, walks, walked, walking); you probably don't want to store all of those forms in your dictionary, instead you'll want to choose a base form from which the others can be derived. Again, tools like FLEx allow you to create morphological parsers that help with that.
Finally, I'd suggest that if you haven't already done so, you should take some courses in linguistics. It will make it much easier to reason about your lexicon, the language's grammar (syntax and morphology), and so forth.
The optimist in me thinks this is the written script for an early version of Proto-Indo-European. Would love to see what like minded people think :)
Meaning a scholar of Sanskrit or something else?
> The optimist in me thinks this is the written script for an early version of Proto-Indo-European
Optimist in what sense?
> Optimist in what sense The one who wants to see a script for PIE discovered :)
I didn't quite understand. Is this related to what the article says: " The discovery of a civilization of people who lived before the Vedic people upended the story of India. Given that it undermines their claims of indigeneity, proponents of Hindutva — the most mainstream strain of Hindu nationalism — balk at the theory of a pre-Vedic civilization, even as evidence for it accumulates across disciplines, including archaeology, genetics, and linguistics. "
Is Sanskrit and Hindutva and PIE related?
I'm just interested in this from a coherence point of view. I believe that learning a super language like PIE would help me understand a lot about human relation to language and vice versa.
Don't know, don't care about this 40 year old concept as much as I do for a 5000 year proto language.
Try deciphering the english language from a bunch of billboards and pop/cans. Even if you figure out the alphabet there aren’t enough words to construct a single sentence.
All the AI in the universe won’t help when the problem is lack of real data.
The goal would be to do that well with languages where we know the grammar and can evaluate the results, and through that find a general tactic. Maybe find a few general theories on how much input that takes, and the character of that input.
If you can extract the grammatical rules with reasonable certainty (not likely with a collection of seals, of course) figuring out the content would just be Mad Libs.
Just training an LSTM deep network on the corpus is enough to do that. GPT-3 is much the same, with a lot more training data and a less computationally intensive network architecture than LSTM, to make that huge amount of training data useful.
Pilus, pirus, pilis. This is intriguing. But one word doesn't signify the membership of a language in another language family. Is it possible that Tamil itself adopted the IVC term for elephants? Of course, since in all likelihood proto-Tamil speakers absorbed many refugees from IVC, there should be many terms shared between the two languages, even if they're from different language families Still interesting to speculate.
Not sure if that's any use here.
The really big problem is lack of data. With sufficient amounts of data and the assumption that we are dealing with human writing, you can probably figure out the encoding.
It is trivial to decode English (or any other alphabetical language) gone through a substitution cypher.
Now, of course, it's not that simple with older languages, but this is a good starting step.
As far as I know, it's only trivial if the substitution is symbol to symbol. If you mess with the letter frequencies, it becomes a lot harder. If you, say, use an orthography that's more in line with the phonology of language in its current form - or devise one that is as quirky, but in a different way. We are talking about a script with 700+ symbols, mathing it together with one that is written in 26 is nonsensical.
This is emphatically not a cryptographic problem. It would really suit ML'ers to learn to take of their shoes and listen a lot more than they speak when they walk into a new scientific discipline.
(BTW, I don't think many modern linguists would accept the term "alphabetic language"; a language is not defined by its writing system, and many written languages are routinely used with multiple writing systems)
Yes it will be more difficult if you play around and change the syntax, join letters, etc. But you usually can figure it out.
> "alphabetic language"; a language is not defined by its writing system
What I mean is a language not written with an abjad, or a syllabary. As in, the actual thing you're trying to decode.
And of course, it may not be a script at all.
As a kid who could read Phoenician script (similar to ancient Hebrew), I tried to decipher the Ashmanezer tablet, which is written in a Semitic language and has many proper nouns. Despite that, there were different ways to interpret the inscription. Taught me just how amazing and difficult language is.
(In the Wikipedia page[1], there is one translation given as though it is empirical. It's not.)
We have no idea what's on those seals. If it's even a script, it might be personal names. The inscriptions are all very short. We have no idea how the script is structured, if it's syllabic or what. We have absolutely no idea what language it's written in.
There is nothing to latch on to, no way to anchor the findings in the real world. The problem with applying AI to this is that it will give an answer, but there's absolutely no way to verify it's a reasonable one.
Encrypted text 'wants' to be obscure, even to those who know the system but lack the key. Most writing systems 'want' to be decipherable to people in the know.
In addition, Enigma has a fairly short key length by modern standards, and lots of accidental flaws that make it even easier to crack.
(Eg an automated system that knows a lot of German _might_ be able to figure out that the enigma never encrypts a letter to itself, even without seeing any schematics.)
The thing is, the written word as an artifact does not itself contain its meaning. To go from writing to meaning is what's called knowing language. The context needed to do that is in the mind of the reader, it's not in the text. Writing something down is quite simply not the same as encrypting it.
What we have here is a few thousand small inscriptions of less than ten characters each. We assume - but don't know for sure - that the characters correspond to sounds, and, if they do, what kind of sounds - syllables, phonemes, maybe a mixture. And even if you could go to sound, then, in order to understand anything, you have to go from sound to meaning, if indeed there is meaning. Which requires knowing the language. And not only do we not know the language, we don't even know what language it is! It might be related to a language of the region, it might not.
Modern cryptography is really good. Without quantum computers, it's unlikely a 16384-bit RSA key will be broken in your lifetime by a motivated actor.
Well, we have a few models of English (large language models like BERT and friends, or the GPT's etc). What insights have they derived about English?
What insights have we derived about English from such models?
Edit: nice paper to read before forming expectations about insights deriveable from language modelling.
Climbing towards NLU: On Meaning, Form and Understanding in the Age of Data.
https://en.wikipedia.org/wiki/Betteridge%27s_law_of_headline...
At one point in college, I was frustrated with an assignment and said “it could either be this or that but I need another datapoint to disambiguate it”, and someone suggested I look at the back of the page.
I've been trying to run away from bad computational linguistics for ten years after making a wrong turn in college. Wish I'd stuck to CxG!
Sometimes I actually think people post these kinds of articles here to ruin my mornings. Did you see the one on the Voynich manuscript where they'd decided it was in Hebrew and done a bunch of maths to it and then, when the person they asked who actually knew Hebrew said "this doesn't make sense", they'd put it into Google Translate and accepted the spell checker's suggestion and then put that as the conclusion of their paper? Fun times.
I like reading as much as the next person, but this kind of articles is starting to get ridiculous.
It's like, it starts with an interesting but otherwise straightforward title, you click, and then you're presented with a novel. One structured in such a way so as to make trying to glance to find the relevant sections yields zero information. Nope, you are required to read all about how someone grew without a dog and how their grandmother had lovely spots on her hands and her skin was like paper, before you can even get an inkling where the information lies in the article, if it's there at all.
I've read countless articles like this by now, where you go through the whole 'novel', and the information you were looking for is on paragraph 67 out of 103, and it's basically the line "Well, we don't really know."
(ノಠ益ಠ)ノ彡┻━┻
Here, let me tell you about the story when I once added 1 + 1 together. You'll never believe the result:"Jonathan was a keen hiker. When his wife got pregnant in the mid nineties, he could never imagine that his great-great-great-great-great grandson would ever be faced with the 1+1 problem [...]"
That's basically the outcome here. It talks briefly about a few different approaches and doesn't really touch on the likelihood of success.
I just grepped for " AI" and read the few paragraphs that popped up. Although looking for "algorithm" proved slightly more interesting:
> For now, the mysteries of the Indus script continue to elude decipherment. Last year, in a follow-up paper to their work automating the decoding of Ugaritic and Linear B, Luo and his team made a small but crucial advance: an algorithm aimed at identifying possible related languages of undeciphered writing systems. Potentially, this could help address the problem of deciphering scripts that don’t yet have a known language they can be compared against. When Luo and his team tested their model on the Iberian language, which has historically been linked to Basque, their findings suggested the two languages were not in fact close enough to be related — a conclusion that corroborated recent scholarship on the matter.
I'd want to see this applied to dolphin clicks. Not sure it would get anywhere, but if an existing human language turns out to have similar grammatical and linguistic constructs, it might get the ball rolling on communicating.
Honestly, the thing I'd want to see this applied to is something like Chinese in Chinese script vs. the same texts in transliteration.