Unicode in identifiers is just a bad idea.
1. It creates a security consideration with confusable identifiers (and lints don't always catch these)
2. It breaks tooling with RTL identifiers
3. It may not render correctly depending on fonts
4. It may be hard to type depending on keyboard layout
5. There really isn't a good reason to use non-ASCII idents anyway
O/0 and I/1/l are confusable characters within ASCII. I'm not kidding here, they are actual entries in the Unicode confusables database [1]. But no one wants to remove those characters from identifiers.
[1] For example, https://util.unicode.org/UnicodeJsps/confusables.jsp?a=0&r=N...
> 2. It breaks tooling with RTL identifiers
It rather unbreaks tooling with no RTL support.
> 3. It may not render correctly depending on fonts
So does Unicode in comments and string literals. In fact the purported Trojan "attack" was mostly about string literals. So why should they be allowed in strings but disallowed in identifiers?
> 4. It may be hard to type depending on keyboard layout
Did you know that not every Latin keyboard layout supports a backquote (`)? This was the actual reason that the repr(expr) shortcut got removed from Python 3 [2].
[2] https://mail.python.org/pipermail/python-ideas/2007-January/...
> 5. There really isn't a good reason to use non-ASCII idents anyway
My canonical answer from the experience is that not every programmer who can understand English documentations can easily write and comprehend English in general. For those people having a non-ASCII identifier support is a great relief, as it frees them from choosing "correct" English identifiers. You can disallow them for your project if you want (or conversely, make it an optional feature disabled by default), but they are relevant for someone else.
Even if you have fluent English skills, sometimes translations just confuse the issue. It's sometimes better to use an untranslated word instead of introducing ambiguity, especially when a term originates from a local law.
Which is why the first thing I make sure of when looking at programming fonts is how well they differentiate these characters
You're mixing up two different ways that people use the word "confusable": things that look similar in some fonts, versus things that look exactly the same regardless of font. I want the latter to be banned from source files but not the former.
[1] https://www.unicode.org/reports/tr39/#Confusable_Detection
Even ones like zero-width joiners and right-to-left marks?
But we are talking about Unicode identifiers, and the Unicode recommendation doesn't allow BiDi markers in identifiers and has a provision to limit the use of ZWJ and ZWNJ in them.
Personally I'd prefer even one step further: the compiler would disallow them by default, and you can opt into specific character sets/languages at a crate level. e.g. `AllowSpecialCharacters("de")` to enable on special characters common in German.
Jokes aside, if you're writing Unicode identifiers it means you're not writing your code to be read by a broad audience.