Yes, we do. People from all over the world write software too. They should be able to use the words they know in code.
Also, it's totally cool to have mathematical symbols in code. λ, for example. Much more readable than the word lambda. The only reason these symbols are hard to type is our keyboards suck. They can be made easy to type with editor support though.
The people working on a non-english codebase don't have to. Their keyboards have the symbols they're typing.
My native language has non-ASCII characters and I do not expect nor do I want to be able to type them outside string literals. Specifically for the reasons stated in the blog post, among others. Writing in my native language is far, far down in the list of priorities as a professional coder, when security / compatibility are there too. Suggesting that non-native English speakers have to be able to code in their native language also would suggest that non-native coders do not take security / compatibility seriously, which would mean that they are unprofessional. I'm pretty sure that it's not your intention to suggest that, but that's kind of how it comes across. With all the problems eliminated by the use of English and ASCII, it would strike me as amateurish to not use English and ASCII wherever possible.
That's not what I said at all. I don't see how you came to this conclusion.
> With all the problems eliminated by the use of English and ASCII, it would strike me as amateurish to not use English and ASCII wherever possible.
Not everybody speaks english. I've taught programming to quite a few people and they all attempted to use normal characters while writing code. There's absolutely no reason why that shouldn't work. I don't see how characters like ç or ã or ü could possibly cause security issues. Go ahead and ban the invisible unicode stuff but there's absolutely no reason why these common letters shouldn't work.
Sure, you could make a fix for this specific case, but the problem mentioned in the blog post is not even close to the only problem of non-ASCII characters. In theory, yes, we could make a language and a full suite of tooling that would play nice with non-ASCII characters. But it's not like the whole non English speaking world is waiting for this to happen. People code in English even in teams where everyone speaks Finnish. Nobody even questions it, because it's so obvious that all code should be in English and ASCII. Everyone has shot their foot, putting in non-ASCII characters in the source code at some point of their career, if they have ever dared to try. That's how the reality is, and at the same time I hear people saying that the existence of those Finnish programmers means we have to have Unicode in source code.
>That's not what I said at all. I don't see how you came to this conclusion.
I didn't say you said it. I said that's how it (probably accidentally) comes across when you talk about something so carelessly. Non English speakers care about compatibility and security and take those seriously, therefore we pretty much always write code in English and ASCII.
Why is it funny? I'm also a member of that group. English is not my native language.
> But it's not like the whole non English speaking world is waiting for this to happen.
I don't think we should have to wait for this to happen. In many ways, it's already happened: most modern languages already support unicode symbols.
> People code in English even in teams where everyone speaks Finnish. Nobody even questions it, because it's so obvious that all code should be in English and ASCII.
Relatively few people speak english in my country. I have only a few friends who do. A whole team of people writing code in english just doesn't seem likely where I live. I actually tried writing english code in such a context once, the result was a mixed language mess that I quickly reverted back to my native language. Unicode support is great because it makes the non-english code much more readable.
Europeans in general seem to know english very well. This is not the case everywhere. Somehow making english a requirement for programming just doesn't sound fair to me.
https://github.com/reinderien/mimic
It applies to other contexts besides code. For our user table we have a mariadb collation on the unicodes confusables list which avoids confusable usernames (treated as already existing).
OK, maybe you're a small startup in Taiwan and so you don't care about the next maintainer in your company not being able to read or write Chinese. What if you decide to open source your code? Or Meta decides to offer you a zillion dollars to buy you out, but after they do their due diligence, realize that the code is utterly unmaintainable should they decide to outsource internationalizing the code so it will work in Brazil, so that requires native Portguese speakers (who can preferably be paid low, low wages) --- but they can't understand the code because it's using Chinese variables and comments. And then Meta decides to back out from the deal?
For example, the school I went to had a simple web application for student feedback. Attachments were allowed. People started running into issues due to non-ASCII characters in file names. I reported the issue to the IT department and even helped them fix it. The Python code was written in portuguese, accents and everything. Why shouldn't accents be used in this case? It's unlikely this code will ever be used in an international context.
It won't. The same approach works just fine in your build specification or other config files. And it doesn't solve the root of this problem, which is that you are compiling source code you don't control and don't audit closely into your binary. Sneaky text is not the only way of getting malicious code through code review.
Jokes aside, if you're writing Unicode identifiers it means you're not writing your code to be read by a broad audience.
Unicode in identifiers is just a bad idea.
1. It creates a security consideration with confusable identifiers (and lints don't always catch these)
2. It breaks tooling with RTL identifiers
3. It may not render correctly depending on fonts
4. It may be hard to type depending on keyboard layout
5. There really isn't a good reason to use non-ASCII idents anyway
O/0 and I/1/l are confusable characters within ASCII. I'm not kidding here, they are actual entries in the Unicode confusables database [1]. But no one wants to remove those characters from identifiers.
[1] For example, https://util.unicode.org/UnicodeJsps/confusables.jsp?a=0&r=N...
> 2. It breaks tooling with RTL identifiers
It rather unbreaks tooling with no RTL support.
> 3. It may not render correctly depending on fonts
So does Unicode in comments and string literals. In fact the purported Trojan "attack" was mostly about string literals. So why should they be allowed in strings but disallowed in identifiers?
> 4. It may be hard to type depending on keyboard layout
Did you know that not every Latin keyboard layout supports a backquote (`)? This was the actual reason that the repr(expr) shortcut got removed from Python 3 [2].
[2] https://mail.python.org/pipermail/python-ideas/2007-January/...
> 5. There really isn't a good reason to use non-ASCII idents anyway
My canonical answer from the experience is that not every programmer who can understand English documentations can easily write and comprehend English in general. For those people having a non-ASCII identifier support is a great relief, as it frees them from choosing "correct" English identifiers. You can disallow them for your project if you want (or conversely, make it an optional feature disabled by default), but they are relevant for someone else.
Even if you have fluent English skills, sometimes translations just confuse the issue. It's sometimes better to use an untranslated word instead of introducing ambiguity, especially when a term originates from a local law.
Which is why the first thing I make sure of when looking at programming fonts is how well they differentiate these characters
You're mixing up two different ways that people use the word "confusable": things that look similar in some fonts, versus things that look exactly the same regardless of font. I want the latter to be banned from source files but not the former.
[1] https://www.unicode.org/reports/tr39/#Confusable_Detection
Even ones like zero-width joiners and right-to-left marks?
But we are talking about Unicode identifiers, and the Unicode recommendation doesn't allow BiDi markers in identifiers and has a provision to limit the use of ZWJ and ZWNJ in them.
Personally I'd prefer even one step further: the compiler would disallow them by default, and you can opt into specific character sets/languages at a crate level. e.g. `AllowSpecialCharacters("de")` to enable on special characters common in German.