Franc seems to work well on longer passages. Such as these: https://github.com/wooorm/franc/blob/master/spec/fixtures.js...
It’s interesting though, I’ll take a look at it!
It’s interesting though, I’ll take a look at it!
I’ll investigate that too. But it’s lots of work, this already was, give me some time :)
And excluding the preamble doesn't make your test more meaningful, the text is still in a very particular domain and writing style.
EDIT: I had assumed trigrams meant word trigrams; character trigrams are a good choice for this.
I haven't it tested on more than 3 languages so it might perform badly but I have the intuition that it is easier to get good coverage of the vocabulary of languages than to get the frequencies of something like the top character n-grams right. The latter is affected by authorship and genre of text &c.