Teach a man to phish and he’s set for life
krebsonsecurity.com
krebsonsecurity.com
Also watch what happens in GNOME desktop environement and its file explorer, like right click on it and view properties, the RLO switch extends to the whole sentence that the filename is written in, it's hilarious.
It's supposed to print "<filename> Properties", but instead the window title is "<beginning of filename>seitreporP <end of filename>".
It's more the development and specs of Unicode that did miss a lot of things. When Unicode came out Bruce Schneier envisioned all the attacks that'd inevitably come and he was, of course, totally right.
As he put it a quarter of century ago (2000): "Unicode is just too complex to ever be secure".
It's Unicode that did fuck up, not developers who are forced to deal with non-sensical things because somewhere a committee decided that it'd make sense to use a broken system for domain names (for example).
Or do you have thoughts on what a "simplified Unicode" — that accomplishes all the same goals — would look like?
Unicode is fine for what it is - conveying text across languages and scripts. The problem is that some environments use the filename as more than what it should be (a memo field for the user and not an instruction for the machine whether to execute the file as code).
What's complex about that? That's just an ordinary sequence of code points, no different than if there was only one language involved.
Unicode has problems where the same thing is represented several different ways. It's not a problem for different things to be represented in different ways.
That's needed for compatibility with legacy encodings.
But when most people — including the GGP poster — complain about this, they aren't usually thinking of canonically-equivalent codepoints, but rather of homoglyphs: codepoints that just happen to have the same conventional pictorial representation. But which have different (collation/splitting/etc) properties; or, more interestingly, different (machine-readable!) semantics.
Most often, phishing "text confusion" is done using homoglyphs. (And, in fact, in some text-processing libraries — e.g. Punycode — "text confusion" can only happen using homoglyphs, as canonically-equivalent Unicode codepoints get normalized to just one option.)
The reason that the homoglyphic Latin "M" and Cyrillic "М" codepoints exist in Unicode isn't for legacy reasons. It's rather because Cyrillic "M" sorts after Cyrillic "Р" (and both of them come before Cyrillic "С"!) You can't have alternate collation orderings if you're using the same codepoints for both languages.
The reason that both Greek "π" and mathematical "𝜋" exist in Unicode, is because one codepoint "is" a letter, and the other codepoint "is" a symbol. This "is"-ness has nothing to do with the glyphs, but rather is basically a type system inside Unicode, that user-agents besides those that render text often rely upon. "𝜋" (the symbol) can carry a particular semantic meaning in text to machine user-agents, that "π" (the letter) does not. Translation systems should treat them differently. Dictionaries should treat them differently. Filter rules for usernames should treat them differently. Etc. (And that's beside the fact that they also have different Unicode properties. "π", as a letter, doesn't create implicit split-boundaries on either side of it; while "𝜋", as a symbol, does.)
However, due to how annoying it is to deal with languages with the same glyphs but a different order ("I don't know where to find things in this translation dictionary!"), languages that share a set of glyphs have tended — ever since democratized access to printing, and dictionaries to create "canonical" orderings — to gravitate toward a shared ordering for the common-denominator subsets of their alphabets.
German, for example, has an alphabet that's like the Latin alphabet but with some "extra" letters — but, however they did it way back when, the German alphabet today "embeds" the Latin alphabet in Latin order at the beginning, and then puts all the extra letters at the end. So German doesn't need another set of Unicode code-points for A-Z; it just needs codepoints for those extra letters.
The extra letters would sort differently though right? I wouldn't expect ö/ä/ü to sort after z. For ß I guess it's not a concern since it never appears at the start of words.
By contrast, GREEK CAPITAL LETTER A is U+0391, separate from LATIN CAPITAL LETTER A. There is no principled reason for this.
> It's Unicode that did fuck up, not developers who are forced to deal with non-sensical things because somewhere a committee decided that it'd make sense to use a broken system for domain names (for example).
I'm not sure that's unicode's fault. Their goal is to produce a standard for encoding ~all written text. It may seem like this would be something cool to support in your program of choice, but in lots of applications that flexibility leads to vulnerabilities. I think developers need to be aware of these risks and design around them (for example by limiting the allowed characters in certain contexts).
> Oregon senator Ron Wyden wants the U.S. government to hold Microsoft responsible for what he describes as “negligent cybersecurity practices” that enabled “a successful Chinese espionage campaign against the United States government.”
https://www.securityweek.com/us-senator-wyden-accuses-micros...
Several people at my org had their email hacked and now I regularly receive my own years-old emails to them sent back to me with phishing links.
Anyone that’s had the pleasure of using office 365 knows microsoft has other priorities.
I’m not even sure this was fixed yet: https://forum.level1techs.com/t/am-i-the-crazy-one-here-or-i...
If anyone is keen to learn more, reach out to me at ksimpson at mailchannels.com. We would like to open source this but don’t have the time to do it properly. If anyone is interested, please reach out to me.
Teach a Nigerian to phish and he'll become a Prince!
Human language for sure was a awesome idea. We're finding out with Generative Large Language Multi-Modal Models that languages are a interesting building block for what we can call intelligence.
Just like with ipv4 and ipv6.
Unicode should have been confined strictly to document content and UI presentation.
Oh but it would make file names less complex. We can't have that.
Edit: btw, how many ways can you express an "ă" in Unicode? How many different file names that look the same visually can you end up with?
This guy expects the automated scanner to render the filename graphically, OCR it, then use the backwards result to guess the file type?
Because first of all, "foo.eml" and "[RLO]bar.eml" are both EML files to any programmatic scanner, and second they look at the content of the file to figure out what it is, not the extension.
I never saw the apology, but even if I had the damage would still be there, the subconcious part of my mind has downgraded ubiquiti
On the flip side maybe ubiquiti deserved to be downgraded - I remember talk about new owners and new directions, and perhaps I never even saw the security claim, and this retraction will unfairly "un-downgrade" them.
It's the same problem as the tabloid press. They can spend 3 weeks hounding someone with lies on page 1, then print a retraction on page 14. Even if the retraction was on page 1 it wouldn't undo the damage.
But I still don't see this as a problem. Allowing arbitrary encodings, arbitrary mix in an single text stream of right-to-left and left-to-right rendering, homonyms and visually identical rendering of different byte sequences are all a problem, and in particular when any of these are being used across a web or OS interface, where recognizing the intent is quite important.
The problem is not Unicode rendering. The problem is improperly mingling information of a different nature. File extensions are not a good way to tell you if a file will be executed (Unix got that right from the start). Then executing a file involuntary shouldn’t lead to significant risks for your data and your system. That’s the actual problem which needs to be solved. Mangling text rendering is not the solution.
That'd be way too extreme but it would have been totally possible to have a filesystem configurable so that it'd enforce ASCII only filenames and ASCII only domain names.
FWIW I enforce ASCII only domain names on my system (I prevent resolving any domain name that has Unicode chars) and live is fine and well. And I'd have zero issue with ASCII-only filenames.
And localization/internationalization (l10n/i18n) can go where it should: inside resource files etc.
In some cases, you can get close enough by just cutting away random symbols. Think é -> e. Sure, you may lose some nuance, but it's close enough. I wonder how well this works for languages which don't use a latin-based script, and how approachable the conversion is for a "regular person".
Not always. For instance, "coco" means coconut, while "cocô" (notice the circumflex, which corresponds to a change on which syllabe has the stress) means fecal matter. That's not just "some nuance".
I love it when people are incensed on my behalf on things I actually don't care about.
No, the word does have a circumflex, not an acute accent, and it's also used by adults (mostly in the expression "fazer cocô", which means to defecate). But, of course, that might be something regional to where I live (a large metropolitan area on the southeast region of Brazil).
You could refuse to render any file encoded in UTF8, UTF16 etc. but that would be silly because by now, most files you're going to encounter (especially as a developer) use one of these encodings.
Or you could refuse to render anything outside of the ASCII range but then you're just assuming that the English-speaking world is all there is and that people of other language communities don't also have needs (such as having their name displayed properly). That's a pretty myopic view of the world IMHO.
I never installed fonts for asian characters on my pc and it hasn't affected me whatsoever. My terminal (which is what I use to manage files) also doesn't support right to left text and it has __never__ been an issue.
Pretty sure you're interacting with lots of people on this website who aren't English native speakers.
> I never installed fonts for asian characters on my pc and it hasn't affected me whatsoever. My terminal (which is what I use to manage files) also doesn't support right to left text and it has __never__ been an issue.
That's not "turning off Unicode", that's just not installing fonts that deal with certain character ranges. Pretty sure whatever font you're using does support more than just bare ASCII, though.
The reasoning is that my primary computer use is in a single language. I very very rarely need to have my computer understand and present to me multiple different languages, and never languages that follow different rules. Anytime such a thing occurs:
a) it is not a language I understand, so not rendering it is fine & safe
b) someone is playing games with Unicode to try to trick me. Definitely not rendering is fine and rendering is actually dangerous
c) the most rare case: the other language is relevant. The only times this has happened (talking to my family and whatnot), it is local to a specific app and I only want that app to deal with the other languages.
The problem is that in the interest of satisfying every language group, my computers assume I care about all languages. I don't, I only speak/read 4 languages and everything else is no use being presented to me. If I'm an outlier, I'm an outlier for being fluent in more than a single language.
For most all users, having the system limited to presenting a single language would significantly reduce the attack surface.
I agree that unicode is ugly and stupid. But despite all that it persists, because it is the least stupid of all encodings. It's not like the ASCII characters of: Form Feed (0C), Device Control 1 through 4 (11 through 14), group separator (1D), and most other non-printable ASCII characters are a reasonable use of byte-space, or are useful to attempt to render to the screen.
And that is before you get into hijinks that can be played with backspace and del characters if they are rendered naively. (Though I love backspace for progress bars in the terminal)
If you just go into simple mode, ë, em-dash, symbols and emoji would be fine.
:^)