Glibc Buffer Overflow in Iconv
openwall.com
openwall.com
In defense of all those usages (except the LSP one, which is indefensible), the original pitches for Unicode[1] literally said that it was intended to be fixed-width, an international ASCII of sorts; that was to be achieved by restricting it to “commercially-relevant” text (and Han unification). Then it turned out there are plenty of very rare Han characters people really, really want to see encoded (for place and personal names, etc.). Of course, in hindsight an encoding with nontrivial string equivalences (e.g. combining diacritics) was never going to be as simple to handle as ASCII.
It's a problem with Unix filenames where the encoding is just a convention. A Python program that doesn't take great care can crash on a parameter that takes a filename even if that filename is just passed to a function like 'open' so no sanitisation or conversion is necessary.
This is a very real problem in hivex, our Windows registry library, where the Python bindings don't really work well. The Windows registry is a hodge podge of random encodings, essentially whatever the program that wrote the registry key thought it was using at the time. When parsed through Python as a string this means you'll get unicode decoding errors all over the place.
Also more in this article: https://changelog.complete.org/archives/9938-the-python-unic...
Don't get me wrong it's still painful and annoying and bug prone, but the point is, it's encodings, it's always going to suck no matter what.
UNIX filenames are NOT necessarily printable text, they're byte strings. Don't treat them as printable text. They're sequences of bytes not containing 0x00 or 0x2F, but with no encoding.
Same for Windows registry keys. Don't mistake byte strings for text.
Text is a byte string with an encoding which describes which byte values are valid and which characters each byte (or sequence of bytes) corresponds to. If the encoding information is discarded, it stops being text and becomes just a byte string.
But they nearly always also happen to be printable text. Treating them as printable text is soooo convenient 99.9% of the time.
Having them be almost printable text, but not quite, is an API design wart that is begging to be used incorrectly. It's the inverse of an affordance.
The POSIX filename rules are a "here be dragons" sign. If you can avoid dealing with them, it's best to do so. If you can't avoid it, you'll need to parse your inputs into valid text, and fall back to a byte string handling path (or just fail with error) if they're not text. You can't safely just assume they're text & treat them as such, it doesn't end well.
(IIRC, at least one of the BSDs has actually moved to forcing filenames to be UTF-8 and refusing path names that aren't UTF-8. Would only that Linux moved down that path as well so that we could be done with this farce.)
I'd agree it'd be nice if we could restrict filenames to valid UTF-8. But that's not the API that existing filesystems provide, nor what (most of) the existing OSes enforce.
Threading both of these needles at once basically requires viewing paths as potentially-invalid encoded text.
How else would you decode a string without knowing its encoding? You can either guess (and risk invalid result/decode errors) or store this information somewhere. This is universal and true in every language. In most cases today people choose to guess utf-8.
>>> 'abc' + 'def'[1]
abce
>>> b'abc' + 'def'[1]
TypeError: can't concat int to bytes
It's been responsible for a lot of bugs in my code. They copied the "bytes is an array of integers" thing from Java. Big mistake. Python is not Java.Them treating filenames as strings is another bug factory on Linux. Almost no one unit tests filenames with invalid encodings, so the result is a whole pile of python3 programs will fall when given perfectly valid input, whereas those same programs in python2 were fine.
It's a very odd outcome given python harks from Linux.
The choice to have this in a language with the safe/ unsafe distinction works very nicely because in so many languages you'd have this promised UTF-8 type and then in practice everybody and their dog uses the unsafe assumed conversion because it's easier, but in Rust you're pulled up short because that conversion needs an unsafe block, your local style may require (and good practice certainly does) that you address this with a safety comment explaining why it's OK and... it just isn't, in most cases. So you write the safe conversion instead unless you really need not to. This is a really nice nudge, you can do the Wrong Thing™, but it's just easier not to.
MFC had a similar problem for years. It had a CString class which was ANSI or Unicode, depending on how your compiled your app, but moderately often you needed the other one, so it should have had a CStringA and CStringW too, with nice conversions between them.
In most of these file format cases what you've got is &[u8] or &[u16] and maybe it's the NonZero variant instead, so I think it's fine to be explicit that's what is going on and maybe in the process remind you to check - is this data UTF16LE? UTF16 with a BOM? UCS2 with a nod and a wink? Just arbitrary 16-bit integers and good luck?
But like I said, I favoured "picking nits" long before I learned Rust, so mileage may vary.
Since references can't have destructors (they don't own the data like an OsString does), it means that the standard library can't give you a newly-allocated string without leaking it. Since obviously it isn't going to do that, the &OsStr must instead just act as a view into the underlying &str. And the conversion can't enforce any extra restrictions on the input string without breaking backward compatibility.
The overall effect is that whatever format OsStr uses, it has to be a superset of UTF-8.
[0] https://doc.rust-lang.org/1.77.2/std/ffi/struct.OsStr.html#i...
That said, even now in 2024, it's not clear how much of a bet Windows is making on UTF-8 versus UTF-16.
Then when 32 bit came to be we had the other variations on top.
Got to love leaky abstractions.
Reference to “Celestial Emporium of Benevolent Knowledge”?
As a meta-joke I was also considering:
APL (j/k)
Just not my cup of tea though.
Rust, C#, Java and Go are fairly straightforward in this regard.
`Encoding.{Name}.GetString/GetBytes`
Either way, I suggest to the readers who might feel upset over this statement to explore something outside of C and C++, liking which, when it comes to strings, is nothing short of Stockholm syndrome.
I'm working on a UTF-8 string library for C# and across the last 6-8 months explored string design in Rust, Swift, Go, C, C++ and a little in other languages. C and C++ were, by far, most horrifying in the amount of footguns as well as the average effort required to perform trivial operations (including transcoding discussed here).
Strings are not easy. But it does not mean their complexity has to be unjustified or unreasonable, which it is in C++ and C (for reasons somewhat different although overlapping). The problem comes from the fact that C and C++ do not enjoy the benefit of the hindsight that Rust had designing its string around being UTF-8 exclusive with special types to express either opaque, ANSI or UTF-16 encodings to deal with situations where UTF-8 won't do.
But I assure you, there will be strong negative correlation here between complaining about string complexity and using Rust, or C#/Java or even Go. Keep in mind that Go's strings are still a poor design that lets you arbitrarily tear code points and foregoes richness and safety of Rust strings. Same, to an extent, applies to C# and Java strings, though they are also safe mostly through a quirk of UTF-16 where you can only ever tear non-BMP code points, which happen infrequently at the edges of substrings or string slices as the offsets are produced by scanning or from known good constants.
If, at your own peril, you still wish to stay with C++, then you may want to look at QString from Qt which is how a decent string type UX should look like.
I wrote more about this here: https://blog.burntsushi.net/bstr/#motivation-based-on-concep...
I mention gecko as an example repository that contains data that isn't valid UTF-8. But it isn't unique. The cpython repository does too. When you make your string type have the invariant that it must be valid UTF-8, you're giving up something when it comes to writing tools that process the contents of arbitrary files.
You sometimes need a way to operate on entirely arbitrary sequences of bytes. This is mostly easy, it's been a long time since non-octet bytes were relevant in most situations, so the vast majority of the time you can just assume they're all octets.
You sometimes need a way to operate on arbitrary text. This inherently requires knowing how that text is encoded, but as long as you know that it's mostly easy.
You sometimes need a way to operate on text-like things that aren't necessarily text, like the output of old CLI programs that used the BEL character to alert the user to events. Or POSIX filenames. Or text where you don't know the encoding. This is where the bugs lie, where we make unchecked assumptions about the data that turn out to be invalid.
You'll notice that I didn't say "Go's string design is good and we should all use it." I made an argument that's Go's string design is not poor and provided an argument for why that is. In particular, I described trade offs and a particular pragmatic point on which abdicating the UTF-8 requirement makes for a more seamless experience when dealing with arbitrary file content.
> but as long as you know that it's mostly easy. [..snip..] Or text where you don't know the encoding.
You don't know. That was my whole point! I gave real-world concrete examples of popular things (Mozilla and CPython repositories) that contain text files that aren't entirely valid UTF-8. They are only mostly valid UTF-8. If I instead treated them as malformed and refused to process them in my command line utilities or libraries, I would get instant bug reports.
> Go strings aren't necessarily text.
I would generally consider this to be an incorrect statement. The more precise statement is that Go strings may contain invalid UTF-8. But the operations defined on strings treat strings as text. For example, if you iterate over the codepoints in a Go string, you'll get U+FFFD for bytes that are invalid UTF-8. By your own reasoning, U+FFFD must be considered text because it can also appear in a Rust &str/String. Despite the fact that a Go string and a []byte can represent arbitrary sequences of bytes, a Go string is not the same thing as a []byte. Aside from mutability and growability, the operations on them (both those provided as a library and those provided by the language definition itself) are what distinguish them. They are what make a `string` text, even when it contains invalid UTF-8.
There are deep trade offs here, but the UTF-8-is-required does have downsides that UTF-8-by-convention does not have. And of course, vice versa.
[1]: https://docs.rs/bstr
https://stackoverflow.com/questions/72143553
> it looks like switching back to UTF-16 would unblock your experience with Resource Editor
Edit: reportedly changed in VS2017 15.9
https://learn.microsoft.com/en-us/windows/apps/design/global...
Getting the hi and lo UTF-16 code points can use the BSD licensed
uint16_t le16toh(uint16_t little_endian_16bits)
and uint16_t htole16(uint16_t host_16bits)They're C99 standard functions and should be converting between "wide strings" and "multibyte strings", which should be native UTF-16 and UTF-8 if your current locale is an UTF-8 locale.
Apparently this works on Windows since Windows 10 version 1803 (April 2018).
There are also "restartable" variants `wcsrtombs` and `mbsrtowcs` where the conversion state is explicitly stored, instead of (presumably) a thread-local variable.
C11 added "secure" variants (with an `_s` suffix) of all these which check the destination buffer size and have different return values.
Well, a conforming implementation could just return -1/EINVAL from `iconv_open()` for any given pairs of character codes.
https://manpages.debian.org/bookworm/manpages-dev/iconv_open...
I wish nothing was in UTF-8 and UTF-8 was relegated to properties files. There are codebases out there with complete i18n and l10n in more languages that most here have ever worked with where there's zero Unicode characters allowed in source code files (with pre-commit hooks preventing committing such source code files).
Bruce Schneier was right all along in 1998 or whatever the date was when he said: "Unicode is too complex to ever be secure".
We've seen countless exploits based on Unicode. The latest (re)posted here on HN was a few days ago: some Unicode parsing but affecting OpenSSL. Why? To allow support for internationalized domain names and/or internationalized emails.
Something that should never have been authorized.
We don't need more of what brings countless security exploits: we need less of it.
Relegated Unicode to translation/properties file, where it belongs.
Sure, Unicode is great for documents, chat, etc.
But everything in UTF-8? emails? domain names? source code? This is madness.
I don't understand how anyone can admire the fact that HANGUL fillers are valid in source code are somehow a great win for our industry.
How about a tonal language? Or a whistling language? Or a clicking language?
I don't, however, understand why you would need support for unicode chars in the source code. To support unicode identifiers? Why would we want that?
I am non-native English speaker. Almost always we work in an international setup, so I already have to reeducate graduates from university not to use localized variable names and not to put in localized comments.
English is the de facto standard in the programs and it's mostly due to technical constraints and historical leadership in programming languages (well, there was French version of Pascal for a while, don't get me started on that ;) ). I think this made us communicate across the borders better. We already have problems of properly naming things, even limited to English, so why would we want to increase the misunderstandings?
This is blatant North American and Western European bias. It makes sense as a policy in those regions, but it is desperately necessary that computing standards do not prevent the rest of the world from participating on as equal terms as possible.
I am Bulgarian, living in Bulgaria, a small country in Eastern Europe. My native language is Bulgarian. I also know English, Russian, and a little bit of French and Turkish.
I do not want to write, or read source code written in Bulgarian, or Russian or French, or Kiswahili, or Japanese etc. I want to use one single language for that. It currently happens to be English for historical reasons.
Supporting unicode in source code, will lead to language fragmentation, and will limit my ability to participate, and to understand everything already build by others.
Please consider that as a (presumably young or young-ish) European, you have access to different educational resources, and are exposed to a cultural context where English is the lingua franca. As a Bulgarian, or any other European nationality, the only foreign language you truly need to learn to get by almost anywhere on the continent is English.
There are many, many countries and places on Earth where English is a significant barrier, due to being much further away, geographically and culturally.
While English is unavoidable after a certain level (just like French is unavoidable for a chef, or Italian is for a musician), it is crucially important that the barrier to entry does not also necessarily come with huge a language barrier up front.
Do they? On my system:
$ grep _ /etc/locale.gen | grep -v UTF-8 | wc -l
183
That's 183 non-UTF-8 locales that are available on my system. OK, I don't have any non-UTF-8 locales currently configured for use, but I don't have to install anything extra for them to be available. Just uncomment some configuration lines and re-run `locale-gen`.https://manpages.debian.org/bookworm/locales/locale-gen.8.en...
I'm struggling to imagine how this failure would manifest. Can you give an example of how dirname() would fail? What combination of existing file/directory name, and usage of that function, would not work as expected?
Edit: I'm also a bit confused how this counts as being a problem for "modern Linux systems" - wouldn't it have always been a problem for all Unix-based OSs?
HTTP theoretically supports Accept-Charset, but it's deprecated:
https://www.rfc-editor.org/rfc/rfc9110.html#name-accept-char...
But I think on-the-fly charset conversion in the web server is quite rare. Apache httpd does not seem to implement it: https://httpd.apache.org/docs/2.4/content-negotiation.html#m...
The charset in question does not have a locale associated with it (it's not even ASCII-transparent), so I don't think it's usable in a local context together with SUID/SGID/AT_SECURE programs.
# from to module cost
alias ISO2022CNEXT// ISO-2022-CN-EXT//
module ISO-2022-CN-EXT// INTERNAL ISO-2022-CN-EXT 1
module INTERNAL ISO-2022-CN-EXT// ISO-2022-CN-EXT 1
Then run "iconvconfig" to rebuild the iconv cache. This disables that charset completely.and it listed ISO-2022-CN-EXT// and ISO2022CNEXT// before I made any changes. After editing the modules and running iconvconfig the command no longer showed those charsets.
This was handy since the alma8 has a /usr/lib64/gconv/gconv-modules file but the file to edit was /usr/lib64/gconv/gconv-modules.d/gconv-modules-extra.conf
I have an old VPS that isn't worth trying to update to a newer OS image, because I'm already (slowly) migrating things off of it before the current paid-up term expires, but it definitely won't get the newer glibc. Disabling the vulnerable character encoding works for me, since no legitimate user of the server will need these conversion pairs.
Wonder what the story is here. Burned 0 day? Not worth exploiting? lolz?
Nothing about this suggests that it wasn't first privately disclosed to the glibc maintainers or that anything else improper happened.