There are plenty of projects out there written by people who aren't English speakers who depend on the Unicode capabilities of languages to write code that is actually readable to them. Turning that off is far from a solution.
There are plenty of projects out there written by people who aren't English speakers who depend on the Unicode capabilities of languages to write code that is actually readable to them. Turning that off is far from a solution.
I myself am not native English speaker and use unicode when writing in my mother tongue, but in 20+ years of programming I've never seen anyone using non-ascii chars in their professionally written code? Of course, you use the language in localization files, and perhaps in comments occasionally - especially in TODO stuff that's not meant to be permanent - but not in the actual code, like e.g. for a variable or function names.
I'd actually consider it a bad idea, as it limits significantly who can manage that code in the future.
How would you name a FooBarWicket if you don't speak a word of English?
I mean don't get me wrong, ideally everybody writes code in perfect English and sticks to a set of ~50 ascii characters, but it's not an ideal world and you have to keep other languages and cultures in mind.
Not sure how would you write a comment in an RTL human language in the middle of LTR code without it. Lots of people write learn RTL languages well before writing any code.
What compilers can do is to process those characters and assign them semantic value that makes the code equivalent to what is expected to be rendered.
Now, bidi overrides in identifier names is a nightmare I’d prefer to avoid.
You only need it if you are doing this, and the default Unicode algorithm for guessing LTR/RTL boundaries gets it wrong, so you need to override with an explicit bidi override control. I'm not even sure how feasible that is to do in current editor/IDE environments developers who have this use case might use.
I am genuinely curious how often these sorts of situations come up in actual development.
> What compilers can do is to process those characters and assign them semantic value that makes the code equivalent to what is expected to be rendered.
I don't understand what you mean or how that's even possible, for the kinds of attacks discussed in OP.
Unicode can handle this, it has a heuristic algorithm for it. Note how if you try to select the text character-by-character, your selection does funny things at the rtl to ltr boundaries, because the byte order doesn't match the order on the screen. It really is handling the directionality changes, with the letters entered in "order" across changes, there is no funny entry or ordering going on, this is plain old normal unicode handling interspersed directionality changes just fine, with no bidi overrides.
It just sometimes gets it wrong for the intent of the author. Especially when there are characters at the boundaries that are themselves not strongly associated as rtl or ltr (like ordinary "western arabic numerals" or punctuation). That's what the bidi override control char is for.
Siht ekil.
It's not even remotely well-defined, and probably never will be. Also, as long as we keep adding to unicode, you will need to keep your whitelist of code points updated.
You can however find a well-defined subset of characters that can be allowed.
In either case you'd be essentially excluding entire languages.
>> There is only ... that should ever be allowed...
What I am saying is someone decides to code in a non-english language (which is completely reasonable) they should define a subset of unicode characters that is acceptable. Additionally, the allowed characters should not permit tricks like these.
As for excluding entire languages... well, yes. This is already the case today. But OTOH it's not like understanding what "if" means gives you any special advantage in programming.
The libraries of most programming languages (developed in the west) are in ASCII - frameworks and middleware too. Have people in countries like Japan and China actually translated all of that code - renaming functions, classes, and variable names to their native tongue in Unicode - or do they just learn the English names (they are all nouns/pronouns and at most simple phrases so translation should not be too difficult; they don’t have to understand English grammar).
China is huge so I can see how it could work for them, but I still have to admit it's very hard for me to imagine someone becoming say a competent web dev without picking at least some basic English along the way, so they can handle at least the documentation and stay in a loop on new tech coming out all the time. It's not anything new as a concept, nor I see it as damaging for local cultures in any way - back in my University days I've learned myself some Russian so that I could read their physics and chemistry books which were excellent and way cheaper and easier for me to get than those from the West. One day I'll have no problem learning some Chinese if (or more likely when?) they become the referent source of knowledge.
Having worked with some large software teams in China my experience was that most people could speak a bit of English (but generally didn't want to) and were nowhere near at the level needed to actually design and write software in English.
If we forced them to do everything in English quality was terrible and everything took ages, but it we let them write in Mandarin things were much better.
Why would they need to learn English to do those things? I'm sure there are Chinese-language tech news sites, and Chinese-language documentation.
We’ll get there eventually with software, but it generally doesn’t kill people so there’s less incentive.
How would you learn how to make a FooBarWicket without knowing a word of English? Any programming languages control constructs are almost by definition English.
(I have used a bidi override before myself, for non-malicious purposes!)
The examples [0] posted in this thread have the bidi characters inside a string literal.
[0] https://gitlab.com/gitlab-org/gitlab/-/commit/3fb44197195b57...
If so, that is devious!
It's not clear to me if those example show that though. They show bidi characters being highlighted in a string literal, right.
My hypothesis was that such could not be part of a "trojan source" attack... but this stuff is confusing and I could have it wrong?
Proper quotes, proper dashes (ASCII doesn't have a dash character, it only has minus), non-breakable space, soft hyphen, € character, Greek letters like π and μ, etc.
Another thing, not every software needs i18n. Depends on the market. I'm yet to see a C++ compiler which would localize their output messages.
Intel C++ compiler seems to have a Japanese version (not tried).
Would you accept teaching code as production code? Specifically, if you were to teach programming to young non English speakers, wouldn't you accept them to use words of their native tongue for variables and such?
> I'd actually consider it a bad idea, as it limits significantly who can manage that code in the future.
Wouldn't you say that solely using roman letters in code would impose a similar limit? In countries where these letters are seldom used (like for instance greek letters in western countries), only those accustomed to them would be able to handle code (as it has been the case until the last decade perhaps).