I’d figured it would be some kind of n-gram frequency analysis. Would be interesting to code that up and compare.
Have you considered doing rune rather than word ngrams? I can imagine that might be prohibitively expensive, but I really don’t know. I did something like that long long ago in C for automatic document language detection. It was quite accurate.