Based on a 2-sec look at the code, it's using a built-in database of trigrams as a predictor of the language.
If the original language data is available I'd suggest classifying the trigrams as "high" and "low" frequency, which should improve performance without needing to keep full frequency data.