I specifically want to address north african languages (Algerian, Moroccan, Tunisian), commonly known as a dialect of Arabic, using a Latin alphabet, with heavy influences from Latin languages (French, Spanish) and Berber (endemic regional language) as well.
My intent is to be able to parse a sentence written in this language, and extract the subject, the action, the object. I would also like to build a semantic model of the language (e.g. word2vec if I am not abusing the field) relating similar words.
I have tried models for Arabic but they fall short because of the borrowings from other languages, but also a different grammar in many cases.
Also, on one hand, transliterating the Latin written sentences in North African Arabic into a regular Arabic alphabet is a challenging task (many transliterations are possible for a given word), on the other, North African Arabic (NAA) is not standardized, so words are commonly written phonetically using a loose set of transcription rules, also borrowed from other languages.
An example of this last aspect would be 'the pharmacist' which in NAA could be written in any of those forms (combinations of either change are possible):
- L'farmassian
- Lfarmassian
- Lpharmassian
- Lpharmacian
- l frmsian
..
Thanks for the reference, I will check it out!