Man if someone asked me to build a system to merely identify whether a unicode string is human language or not I would flatly refuse. There are thousands of spoken languages, many of them with no standard written form, some that are transcribed into multiple different writing systems, some with no writing tradition at all and with only ad-hoc transliteration unique to each user and use.
Even being 90% confident would be a massive undertaking, and "speakers of this language may/may not use the internet" feels like high stakes for getting it wrong.
It seems a little niche but I'm sure a few times a year some far out town gets connected and suddenly there are speakers of a previously unknown-to-the-internet language newly online.