Show HN: I trained a neural network to learn Arabic morphology
github.com
github.com
This is an active area of research in Morphologically Rich Languages (MRLs), since this problem also appears in other semitic languages like Hebrew, as well as Turkish. There's a nice body of work to learn from, both with and without neural nets. For example, this paper from 2017 (http://aclweb.org/anthology/D17-1073) uses a neural model for morphological disambiguation. You can see a nice comparison of tools in the recent 2018 Universal Dependencies Shared Task results: http://universaldependencies.org/conll18/results-lemmas.html (look for ar_padt).
If you're looking for training data, the Arabic treebanks in http://universaldependencies.org could help. I think some of them contain surface tokens with lemmas. I'm quite sure they also have roots.
Also, you might want to take a look at the SIGMORPHON CONLL shared task (2017 https://sites.google.com/view/conll-sigmorphon2017/ and 2018 https://sigmorphon.github.io/sharedtasks/2018/) on morphological reinflection, which IIRC is a similar task - taking an inflected form and reinflecting it with other morphological properties. They also have a nice data set to train on.
There should be ones specializing in Arabic or morphology, but the generic LREC is a quite wide one, e.g. http://lrec2018.lrec-conf.org/en/conference-programme/accept... has papers that seem relevant like "Build Fast and Accurate Lemmatization for Arabic" (http://www.lrec-conf.org/proceedings/lrec2018/pdf/1079.pdf), "Part-of-Speech Tagging for Arabic Gulf Dialect Using Bi-LSTM" (http://www.lrec-conf.org/proceedings/lrec2018/pdf/483.pdf), or "A Morphologically Annotated Corpus of Emirati Arabic" (http://www.lrec-conf.org/proceedings/lrec2018/pdf/529.pdf).
Along those lines, I might be able to provide some useful json (the basis for http://www.arabicreference.com) in case you are interested. I've been meaning to do some fun investigations using this data (e.g. predicting broken plurals, masadir, form I internal vowelling) but haven't yet had time.
1. http://oldsite.bayyinah.com/wp-content/uploads/2015/12/sarfG...
I know nothing about arabic, so to my eyes certain wiggles - that aren't a 1:1 match - are "correctly" matched, and others aren't.
ex. Correct. Input: احتياطي (AHtyATy). Answer: حوط (HwT).
So what the NN does is takes as input a word, and tries to find its root. "Tabdeel" was the first input listed, and the output was "BDL".
Some more information on this:
So the root d-r-s has derived verbs darasa and darrasa, and to each of those correspond, say, one or more patterns for the verbal noun. But I don't think there is exactly one pattern for verbal nouns derived from the form 2 verb (e.g. from darrasa we get tadris, as I recall, but not all verbs that go like fa33ala will necessarily have a masdar of the form taf3il, right?).
You're right of course, that even though the forms have prototypical systematic semantic variation (like form 2 is usually a causative, "to teach" versus "to learn"), it's not predictable which derived forms of a given root enter into actual usage and with which exact meaning, and the patterns obviously predict a lot of words that don't actually exist, and of course Arabic speakers learn words just like speakers of any other language.
I think I remember that there are a handful of cases where speakers started using some of the previously unattested forms in modern times to refer to new concepts... is that right?
(What a great topic - I miss this stuff!)
How big is the problem space? There's a limited set of roots and morphologies...
Would a more rules based approach work (more accurately)?
For the purpose of just using the language, the morphological rules are well-understood. One of the most popular dictionaries (Hans Wehr) is arranged not in alphabetical order, but by root. And there are many online lexicons with morphological metadata as listed in other comments here. So you're right, machine learning is not necessary here (but that's not to say it couldn't be done). This is mainly just for fun and learning.
Now, if you were to add all dialects of Arabic, then you may have a use case for ML...
Edit: thanks for the response, it is interesting result that more layers don't improve the situation.
جملة means a "sentence"
جمل can mean a Camel (pronounced "jamel" you can guess the origin of the english word)
جمل can also mean "sentences" (pronounced "jomal")
جَمَلَ means to sum-up, to summarize[1] or to concatenate. Which is what a sentence does to words.
You can write out the short vowels but it's only done in special contexts (like children's books, books for foreign language learners, or the Qur'an). It's easy enough to read if you know the language, but it makes it a little harder to learn.