A Tale of Two Libcs
drewdevault.com
drewdevault.com
When you encounter other peoples' code and it looks a bit complicated, and not the 'obvious and simple' thing, your first thought should never be 'they wrote bullshit code', but instead 'what have I missed? What did I forget to consider for this function?'
I don't know anything about the glibc authors, or their code style, but if you go through your life assuming incompetence if someone wrote a function differently to your initial thoughts, then you are an arrogant prick.
But if there's something I'm sick of it's that stupid "well the spec says it's UB so let's just write code that FUCKS THE DEV as hard as possible." attitude in certain circles regarding a one letter programming language.
The code in question is a classic array lookup with missing bounds check. If you write code like that and then blame people for using it wrong because the spec says so, you're just as arrogant, just in a different way.
Tangent: surprises me how often comment people lose their shit when you utter the slightest word of criticism about some pet {language,code, tech}, but are absolutely fine, supportive and enabling when someone verbally eviscerates the living shit out of another human based on what they said, or their opinion.
How is attacking a non human idea and no one in particular bad, but attacking a specific person viciously ok? I don't get that and I never did.
Well it's not an absolute, is it? There's context. Some specific people deserve to be attacked.
The function that they were calling can be summarized as, "check to see if a value is in a set," only to receive an initially inexplicable segfault. They had a clear example of what that could look like, someone else's code from a common library, yet spent a considerable amount of time untangling the code in the library where the bug popped up.
It is wrong to assume incompetence on the part of the glibc authors, particularly when the comparison library has a much more limited scope and can afford to avoid "bullshit code", but the source of the frustration is clear. I would hardly define them as an arrogant prick because of that.
My opinion is that there is only one excuse for writing so unreadable code for such simple functionality (even taking account the legacy): that would be to use a language completely unfit for the task. That should not be the case here ... or is it? Why so many macros?
I've written tons and tons of software for Linux, and I believe only once have I seen a bug in glibc[1].
[1] And it was both a horrible bug, and also an extremely weird corner case: https://bugzilla.redhat.com/show_bug.cgi?id=563103#c8
The answer is a pretty big yes. Glibc's isalnum is about 6-7 times faster than musl's.
Benchmark code: https://paste.sr.ht/~minus/18b44cfe58789bc1fb69494130e859a11...
For those interested: The technique is used in P. J. Plauger's 1992 The Standard C Library and has been around a long time.
And even if you want to use a LUT, other implementations have incorporated a LUT without this segfault issue, for instance plan 9:
https://code.9front.org/hg/plan9front/file/00397858f3d8/sys/...
https://code.9front.org/hg/plan9front/file/00397858f3d8/sys/...
And I fully agree with you that this glibc code is awful. As evidenced by the Plan 9 implementation, table lookups can be done simple as well, and probably perform just the same.
>So the fix is obvious at this point. Okay, fine, my bad. My code is wrong. I apparently cannot just hand a UCS-32 codepoint to isalnum and expect it to tell me if it’s between 0x30-0x39, 0x41-0x5A, or 0x61-0x7A.
If that's what you want, just go ahead and write that.
I understand both sides of the argument here - on one hand you have fast and unreadable code and on the other hand slower and more readable code. The problem here is 1. the leaky abstraction of characters and code points and 2. loss of type safety due to C's weak typing rules, which is a bit lost in the "hurr durr messy code."
Do you have a source for this claim?
I find it hard to believe that Microsoft would drop support for ANSI applications.
- Support for weird compilers
- Focus on performance, thus playing "weird" tricks with macros, lookup tables etc.
- especially in headers: protect against users doing crazy sh*t
- standards sometimes change or see extensions
- certainly more :)
First two items can be re-evaluated form time to time, but especially GNU often aims to run "anywhere"
The third item comes from the fact that users can create macros for many things and only some names (leading _ and capital letter or __) are reserved for standards or C++ allowing to overload the comma operator, which can have weird effects. Thus to be conformant to the standards some ugliness is required.
The fourth point combined with different compatibility requirements requires tricks to switch between different implementations based on compiler flags/macros (for instance `#define _GNU_SOURCE` or `#define _POSIX_C_SOURCE=...`)
> an unsigned char or […] EOF
so isalnum is "their own specific variants of character related functions" for uchar, and locale-aware.
This is bullshit. Great, now isalnum returns true for ä if I'm using de_DE.latin1 or whatever, but is still nonsense for any non-western language. And the price we pay for that is a messed up pile of garbage that nobody understands and can segfault.
But all the platforms that use glibc have moved to utf8 over a decade ago. So even for western languages that code doesn't make sense anymore.
I'd argue the usefulness of all that glibc code is very close to zero. Sure, the glibc folks can pat themselves on the back for covering the POSIX spec so well but I really prefer the Linux approach here; follow POSIX where it makes sense but omit the insanity.
Your comparison to Linux makes no sense either. It's not that the code pages themselves are the problem, you can use them for conversion and whatnot. But pretending a collection of functions that is unable to do anything meaningful in the present day and crashes for extra points is fine, because it was written long ago just isn't. The musl solution is correct. If you need to handle strings in a locale-aware manner, use a dedicated lib. Better yet don't use C.
And the crash is a problem with C but it's not a problem with the function. There's a different function, wisalnum, that doesn't crash with out of bounds values, but it also doesn't do Unicode on musl.
The crash is a problem with the function, specifically with glibc. You might even put into question why POSIX defines out of range values as UB instead of requiring it to return false but that's yet another can of worms we better not open here. I fully blame not doing any range checks and doing an oob array access on glibc. UB doesn't force you to create an implementation that crash and burns... Returning false would be great. Maybe abort if you're anal. Have it return something random if you're concerned about speed. But don't freaking crash! The libc should work with the dev, not against them.
It doesn't.
> This is bullshit.
Welcome to POSIX locales.
> Great, now isalnum returns true for ä if I'm using de_DE.latin1 or whatever, but is still nonsense for any non-western language.
Technically it's already nonsense for western languages as ISO-8859 is long outdated, and any codepage other than ISO-8859-1 will not fit inside a uchar when decoded from UTF-8. Also there were non-western languages which fit in 8-bit encodings for which isalnum would work fine.
> And the price we pay for that is a messed up pile of garbage that nobody understands and can segfault.
Welcome to POSIX locales. And also C, where not reading and understanding the implication of every word in the specification means you're bad and therefore deserve everything you get. In this case, POSIX clearly specifies that input values outside of EOF or "unsigned char" is UB.
It probably just shows how old the glibc is and like most software entropy is increasing over time.
Musl assumes your locale is either C or UTF-8, glibc doesn't. Sure almost everything is UTF-8 these days and has been in the last 15 years, but claiming that the glibc code makes no sense is utterly condescending and I expected ddevault to know better than that.
I also wouldn't be surprised if those countries that have a non-Latin alphabet were still using single byte character encodings since UTF-8 would double the size of their text.
That's more complicated because depending on the language these may or may not be considered separate letter e.g. the french alphabet has 26 letters: while required by the grammar, diacritics and ligatures are "extras". This means isalnum wouldn't necessarily return true for, say, ç in the french locale (I've no idea what it actually does).
Turkish however does have 29 letters, and the dotted / dotless i is very much part of the alphabet (so are ç, ş, ğ, ö and ü).
> Musl assumes your locale is either C or UTF-8
UTF-8 is an encoding, not a locale. And Musl just plain assumes the C locale, it clearly has no support whatsover for anything else.
Which, for what that's worth, is a respectable choice: posix locales are really really bad, so not bothering with them is not necessarily a negative, but at the same time dinging glibc for supporting locales or old platform is not very honest.
It returns true, otherwise "advance to the next word" or "grep -w" would be utterly broken. ("grep -w ça" would not match anything, and in fact it probably won't match anything under musl).
Other European languages also treat characters with diacritics as separate letters, for example the Czech alphabet has 42 letters (one of which is "ch" which doesn't have its own Unicode character as far as I remember; if it did there would be more fun to be had with normalization and title case).
> And Musl just plain assumes the C locale, it clearly has no support whatsover for anything else.
Yeah what I meant is that if you are reading UTF-8 you won't be passing chars in the 128-255 range to the ctype functions. But musl doesn't implement Unicode wchar_t either.
The trap is that ctype functions take int, that's exactly the mistake TFAA made: they passed in an integer assuming isalnum would work with codepoints (which doesn't actually make sense since int is only required to be 16 bits).
It looks like it does. iswalpha uses some kind of table to return its response. Is that implementation somehow too simplistic?
The musl isalnum code is completely locale-unaware. Which you might argue is a good thing as POSIX locales are awful, but still.