I think I found it in this paper [1]. The implementation was like 13 lines of Python code. I wonder how it would compare.
[1] http://www.ccs.neu.edu/home/jaa/CSG399.05F/Topics/Papers/Ben...
Also, I’m interest in a test-suite, before we start talking about accuracy-percentages :p
Although I don't need a test suite to confidently say that L1 distance is going to be worse than naive bayes, it would indeed be good if you had a test suite!
We already have a test suite with 1 example: https://news.ycombinator.com/item?id=8405180
Naive bayes would never get something like that wrong.
It’s interesting though, I’ll take a look at it!
I’ll investigate that too. But it’s lots of work, this already was, give me some time :)
And excluding the preamble doesn't make your test more meaningful, the text is still in a very particular domain and writing style.
EDIT: I had assumed trigrams meant word trigrams; character trigrams are a good choice for this.
I haven't it tested on more than 3 languages so it might perform badly but I have the intuition that it is easier to get good coverage of the vocabulary of languages than to get the frequencies of something like the top character n-grams right. The latter is affected by authorship and genre of text &c.