Show HN: My Leeds Hack Day project, a language detection API
polyglossy.com
polyglossy.com
It uses a slightly weird technique, though. Dictionary based and using a bloom filter for memory efficiency. Going forward, though, I plan to rewrite it to use a combination of n-grams and language "fingerprints."
Do you have any accuracy stats? I'm guessing my approach might work better in some cases because the models include frequency information too. Did you experience significant accuracy loss when adding new languages? Anyway, I'll run it over my test data and compare.
No, but as you have noted, the method has the intrinsic property of being less accurate with fewer words and more accurate the longer the text. As my anticipated use was for documents over 10-20 words, this was OK. I expect the other techniques I outlined that I'm switching to to yield more accurate results across the board.
How does this work? You claim it has "zero language knowledge," so how does it classify? Did it start out with nothing, and then train it on a corpus?
EDIT: nevermind, found your slides (http://polyglossy.com/presentation.pdf [pdf]) How big of a corpus did you use?