HNHacker News
TopNewBestAskShowJobs

wooorm

44 karma · joined July 31, 2014

submissionscomments
wooorm··on Show HN: Alex – Catch insensitive, inconsiderate writing
alex is open to suggestion. See http://alexjs.com/#contributing on how to contribute. So if you have a word to add or phrasing to remove let us know!
wooorm··on Show HN: Franc – Detect natural languages
Perhaps you should ;) If, I’d be interest to know how it goes!
wooorm··on Show HN: Franc – Detect natural languages
It’s a very interesting idea. Would it work accurate enough when scaled to 160+ languages?
wooorm··on Show HN: Franc – Detect natural languages
I pushed a fix, incorporating your suggestions, and your examples in the specs.

Thanks a lot!

wooorm··on Show HN: Franc – Detect natural languages
It’s an interesting thought. I might fiddle on it, but I’m not sure it would work in practice (d’oh). Thanks!
wooorm··on Show HN: Franc – Detect natural languages
I agree the task is neither impossible nor useless. There’s work to do. Short passages should be supported. I do however think franc does a good job, and adds support for some languages which before today have never (I think) been supported. Franc, certainly, “attempt”s to fix language detection, which I would argue is an AI-complete problem.
wooorm··on Show HN: Franc – Detect natural languages
Thanks ;)
wooorm··on Show HN: Franc – Detect natural languages
By `correct language` I mean the language you expect, by `second` and `third` I mean `2.` and `3.` in the previously mentioned demo: http://wooorm.github.io/franc/). I think we’re talking about the same thing!

Anyway, yeah, franc is for language detecting, but it’s optimised for many languages and works best at longer text. It’s a trade-off. For less languages and shorter texts, check out https://github.com/shuyo/ldig

wooorm··on Show HN: Franc – Detect natural languages
Ha! Some very nice examples, I have to say :)

Anyway, You’re completely right. Italian is `und` due to LTE 10 characters, the others are slightly off due to short input too, but the demo (http://wooorm.github.io/franc/) shows the correct languages in the second or third place though!

wooorm··on Show HN: Franc – Detect natural languages
No full-frequency data is kept, only 300 top-trigrams are identified. A quick through the source also reveals wooorm/trigrams, and wooorm/udhr, as sources!
wooorm··on Show HN: Franc – Detect natural languages
Oh you’re right. I think I have a fix in mind, will work on it. Thanks so much!
wooorm··on Show HN: Franc – Detect natural languages
I’ll investigate this, but I think I excluded the preamble’s for trigram creation. Sure, the words will be a bit similar, but it’ll be a lot of work to compile 380 fixtures from other sources.

I’ll investigate that too. But it’s lots of work, this already was, give me some time :)

wooorm··on Show HN: Franc – Detect natural languages
Thanks! Currently, the UDHRs are crawled, and I’d rather not include exceptions and maintain their plain-text and XML/JSON versions by hand. If you’re into growing the language, I suggest contacting the Office of the High Commissioner of Human Rights of the UN, and the Unicode project, or fork wooorm/udhr and add support, I’ll merge :)
wooorm··on Show HN: Franc – Detect natural languages
Franc seems to work well on longer passages. Such as these: https://github.com/wooorm/franc/blob/master/spec/fixtures.js...

It’s interesting though, I’ll take a look at it!

wooorm··on Show HN: Franc – Detect natural languages
It sucks, right? Currently, it’s good at long passages. But for shorter values, the results are pretty poor. The amount of supported languages is just too damn high!
wooorm··on Show HN: Franc – Detect natural languages
That would be awesome :)
wooorm··on Show HN: Franc – Detect natural languages
I’m not sure. I don’t know any CJK languages myself. I’d like some test-cases where the current methods do not work, as the example in the Readme seems to work pretty well: `এটি একটি ভাষা একক IBM স্ক্রিপ্ট` is classified as Bengali?
wooorm··on Show HN: Franc – Detect natural languages
And it doesn’t have a Universal Declaration of Human rights: http://www.unicode.org/udhr/index_by_name.html
wooorm··on Show HN: Franc – Detect natural languages
Fries as in Frisian? I don’t think it has one million speakers (right?) :p
wooorm··on Show HN: Franc – Detect natural languages
You seem to be completely right, I hand-crawled the data (https://github.com/wooorm/speakers), but seem to have made big typo there! Thanks!
wooorm··on Show HN: Franc – Detect natural languages
One of franc’s focusses was to be pretty small, and usable on the client-side, that’s why no actual training is done and this simple method is used.

Also, I’m interest in a test-suite, before we start talking about accuracy-percentages :p

wooorm··on Show HN: Franc – Detect natural languages
Yeah, so I’d like to add an easier way to support more, or less, languages through the Node API. Currently, there’s a number (1e6), the amount of speakers of a given language, which is hard-coded in the generation file (I added a link this morning in the statement about forking to the actual line).

If you set that number to 0, or 100,000 and execute `npm prepublish`, your franc supports more languages :) That’s it!

wooorm··on Show HN: Franc – Detect natural languages
Agreed :)
wooorm··on Show HN: Franc – Detect natural languages
I’m also really interested in trying something like this: http://www.slideshare.net/shuyo/short-text-language-detectio... (slide 6). But I’d need a lot of training data, more than UDHR.
wooorm··on Show HN: Franc – Detect natural languages
That’s because Haitians always say that! No, joking, it’s just that because of so may supported languages, the accuracy for very short inputs is extremely low.
wooorm··on Show HN: Franc – Detect natural languages
You are completely right, franc doesn’t state how language are detected. The detection is based on (1) unicode-script usage and (2) trigram-counts. Some scripts are only used by one language. Other scripts, such as Cyrillic, come with many more: those are detected by the top 300 trigrams per their corresponding UDHR (Universal Declaration of Human rights, the most translated document).

Shameless plug, the page clearly stated you can fork franc to support 300+ languages ;)

wooorm··on Extensible system for analysing and manipulating natural language
Hey all, the creator of Retext here! Please let me know if you have any recommendations, issues, or general feedback :)