I think that depends on the users. People copying and pasting bits of text that was in English or another common languageโ think documentation, code, news articles, tweets, etc.โ with a different character set could be problematic.
Also, ๐ฎโด๐โฏ ๐๐ ๐ ๐ marketed as "๐ฝ๐ ๐๐ฅ๐ค ๐๐ ๐ฃ ๐ค๐ ๐๐๐ ๐๐๐๐๐" would be โญ๐๐ฒ๐ค๐ฅ๐ฑ ๐ฒ๐ญ ๐ฆ๐ซ ๐ฑ๐ฅ๐ฆ๐ฐ. (math symbols) A user base with young people getting bounced or shadow banned for trying to express themselves or distinguish themselves from their peers would be like เฒ _เฒ (Kannada letter ttha)
I think targeting the language they're using is a better bet.
ยฏ\_(ใ)_/ยฏ (Hirigana letter tsu)
Minor nitpick, but ใ is the katakana tsu.
If they use many (maybe three? four? or more) character sets in the same post, or different character sets in any single word, then that'd be highly suspicious?
Whilst still letting people copy paste from another language
Special case needed for the shoulder shrug with an Hirigana letter tsu I mean katakana tsu
Sounds like you live in a filter bubble.
(โฏยฐโกยฐ)โฏ๏ธต โปโโป
(๏พโใฎโ)๏พ*:๏ฝฅ๏พ
They kinda do. Check out the shrug "emoji", table flip, and so forth. Then there's the meme of adding text above and below by abusing Unicode's "super" and "sub" modifications.
You could block it to only ever represent ASCII, but then you've knocked out the ability to expand internationally.