Show HN: Franc – Detect natural languages
github.com
github.com
I think I found it in this paper [1]. The implementation was like 13 lines of Python code. I wonder how it would compare.
[1] http://www.ccs.neu.edu/home/jaa/CSG399.05F/Topics/Papers/Ben...
Also, I’m interest in a test-suite, before we start talking about accuracy-percentages :p
Although I don't need a test suite to confidently say that L1 distance is going to be worse than naive bayes, it would indeed be good if you had a test suite!
We already have a test suite with 1 example: https://news.ycombinator.com/item?id=8405180
Naive bayes would never get something like that wrong.
It’s interesting though, I’ll take a look at it!
I’ll investigate that too. But it’s lots of work, this already was, give me some time :)
And excluding the preamble doesn't make your test more meaningful, the text is still in a very particular domain and writing style.
EDIT: I had assumed trigrams meant word trigrams; character trigrams are a good choice for this.
I haven't it tested on more than 3 languages so it might perform badly but I have the intuition that it is easier to get good coverage of the vocabulary of languages than to get the frequencies of something like the top character n-grams right. The latter is affected by authorship and genre of text &c.
shameless plug https://github.com/allan-simon/Tatodetect it covers 179 language (actually as much as Tatoeba project does) and it can run offline with explanation on how to generate your own database from a CC-by corpus.
After the advantage of Franc is that it can be used directly as a npm library while Tatodetect is a micro-webservice, and for some edge languages, Tatodetect is certainly not as good as Franc (haven't done yet a test of both to compare)
Shameless plug, the page clearly stated you can fork franc to support 300+ languages ;)
If you set that number to 0, or 100,000 and execute `npm prepublish`, your franc supports more languages :) That’s it!
If the original language data is available I'd suggest classifying the trigrams as "high" and "low" frequency, which should improve performance without needing to keep full frequency data.
That's quite common, like in mixed-language IRC channels, quotes from English documents in documents mostly written in another language, and so on.
And stemming and indexing such a document for full text search usually gives crappy results.
(Bonus points of detecting programming code samples, so that this part isn't stemmed at all).
Sometimes it gets it almost right: I tried with this piece of text in Catalan (Balear variant) and it classifies it as Portuguese (with Catalan as 2nd option): "I s'horabaixa la deixam passar i me mires tan a prop que me fa mal, que surt es sol i encara plou, que t'estim massa i massa poc, que no sé com ho hem d'arreglar, que som amics, que som amants."
It's strange, because it's pretty different from Portuguese...
The Catalan poem "tirallonga de monosíl·labs" gets classified as French. (http://www.rodamots.com/calaix.asp?text=tirallonga)
CJK scripts and languages tend to be relatively more concise (in terms of # of Unicode codepoints) than many other languages, so it is possible that the ratio of CJK scripts over non-CJK scripts can be lower than the average. And the occurrence ratio is currently calculated over the number of characters including non-letters, making the ratio much lower. Maybe the custom threshold per script based on the actual corpus (90th percentile, maybe?) and better occurrence calculation would improve the detection on those languages.
한국어 문서가 전 세계 웹에서 차지하는 비중은 2004년에 4.1%로, 이는 영어(35.8%), 중국어(14.1%), 일본어(9.6%), 스페인어(9%), 독일어(7%)에 이어 전 세계 6위이다. 한글 문서와 한국어 문서를 같은 것으로 볼 때, 웹상에서의 한국어 사용 인구는 전 세계 69억여 명의 인구 중 약 1%에 해당한다.
This text from Korean Wikipedia is about the ratio of Korean documents over all documents in the Internet. Digits distort the overall ratio and Franc doesn't give any candidates (even no "und").
現行の学校文法では、英語にあるような「目的語」「補語」などの成分はないとする。英語文法では "I read a book." の "a book" はSVO文型の一部をなす目的語であり、また、"I go to the library." の "the library" は前置詞とともに付け加えられた修飾語と考えられる。
This text from Japanese Wikipedia concerns about the distinction of objectives and complements in the English syntax. In this bilingual text it looks like that Japanese has reached the 60% threshold but the codepoint count doesn't.
Thanks a lot!
ron? snn
fra? cat
swe? nds
ita? und
nld? gax
Source: var franc = require('franc');
console.log('ron?', franc('Cate bere ai baut?'));
console.log('fra?', franc('C\'est quoi le bordel la, putain'));
console.log('swe?', franc('Jag kanner en bot, hon heter Anna'));
console.log('ita?', franc('che guai'));
console.log('nld?', franc('graag gedaan'));https://en.wikipedia.org/wiki/List_of_English_words_of_Frenc...
…do you expect it to return French or English?
Yes you can present edge cases where there is no definite answer, like the one you cite, but this doesn't mean that the task in general is impossible or useless.
Anyway, You’re completely right. Italian is `und` due to LTE 10 characters, the others are slightly off due to short input too, but the demo (http://wooorm.github.io/franc/) shows the correct languages in the second or third place though!
Anyway, yeah, franc is for language detecting, but it’s optimised for many languages and works best at longer text. It’s a trade-off. For less languages and shorter texts, check out https://github.com/shuyo/ldig
Though I'm sure your test wasn't intended to be insidiously misleading.
after you can also try to explan them that the common "represent a language by a flag" becomes quickly broken and subject to strong arguing between people (what flag do you put for Tibetan language for example? or for each of Indian languages)
you have a CSV of iso code => sentence , which should be 99% accurate (as it gets user proofed), so on in which you can compare your tool with.
I think for longer text one could use Wikipedia dump or alike ?
Unfortunately Fries is not supported, but I'd be interested in the results. But I don't think polyglots for natural languages are common, this is in fact the only one I know.
P.S. Kudos, very cool project!
EDIT: Frisian version should you want it: https://www.google.com/search?q=Yn+betinken+nommen+dat+it+er...