Looking into generated ngrams, I’m not sure that it’s good idea of getting rid of spaces between words, and not having markers for begin/end of word.
It would be interesting to check on the dataset linked to this blog post from 6 years ago, evaluating fasttext model for language detection: https://alexott.blogspot.com/2017/10/evaluating-fasttexts-mo...