https://github.com/wanderingstan/Confusables
E.g. "𝓗℮𝐥1೦" would match "Hello"
Maybe somebody else on here will see it and learn about it before they need it, and at least you still have a new tool to reach for in the future.
>>> import unicodedata
>>> unicodedata.normalize('NFKD', '𝓗℮𝐥1೦𝗵𝗲𝗹𝗹𝗼')
'H℮l1೦hello'
Obviously it isn't going to remap leetspeak characters like 1 -> l but it covers a lot of cases.NFKC/NFKD will handle "this is another form of the Latin letter A" type stuff but not "Cyrillic A looks like Latin A."
Back in 2015 I did some work just using simple bitmap rendering plus OCR to find text that looked like a small selection of known words. It was actually reasonably effective.
Unicode Character “;” (U+037E) is a beaut.
The initial problem wasn't those symbols but the content itself, the symbols and special characters came into the problem later.
Later on as mentioned in my original comment, that they would use positive content from other blog posts that were published/passed the moderation to mix up their bad content.
Probably could use a different method, but at that time needed something quick and fast and it worked and still works with very little tweaking.
Although we don't have massive amount of threats or abusers anymore to exactly know the effect, but again, so far it works.
That time, they would coming several thousands per minute, IP blocking, range blocking, USER AGENT, captcha or anything such didn't work on them.
I guess I can use that next time time to work on the data cleaning for that model.
Thanks.
I built a Python library for finding strings obfuscated this way. Was critical when moderating our telegram channel before an ICO. https://github.com/wanderingstan/Confusables E.g. "𝓗℮𝐥1೦" would match "Hello"
They kinda do. Check out the shrug "emoji", table flip, and so forth. Then there's the meme of adding text above and below by abusing Unicode's "super" and "sub" modifications.
You could block it to only ever represent ASCII, but then you've knocked out the ability to expand internationally.
I think that depends on the users. People copying and pasting bits of text that was in English or another common language— think documentation, code, news articles, tweets, etc.— with a different character set could be problematic.
Also, 𝒮ℴ𝓂ℯ 𝒜𝓅𝓅𝓈 marketed as "𝔽𝕠𝕟𝕥𝕤 𝕗𝕠𝕣 𝕤𝕠𝕔𝕒𝕝 𝕞𝕖𝕕𝕚𝕒" would be ℭ𝔞𝔲𝔤𝔥𝔱 𝔲𝔭 𝔦𝔫 𝔱𝔥𝔦𝔰. (math symbols) A user base with young people getting bounced or shadow banned for trying to express themselves or distinguish themselves from their peers would be like ಠ_ಠ (Kannada letter ttha)
I think targeting the language they're using is a better bet.
¯\_(ツ)_/¯ (Hirigana letter tsu)
Minor nitpick, but ツ is the katakana tsu.
If they use many (maybe three? four? or more) character sets in the same post, or different character sets in any single word, then that'd be highly suspicious?
Whilst still letting people copy paste from another language
Special case needed for the shoulder shrug with an Hirigana letter tsu I mean katakana tsu
Sounds like you live in a filter bubble.
(╯°□°)╯︵ ┻━┻
(ノ◕ヮ◕)ノ*:・゚