I understand both sides of the argument here - on one hand you have fast and unreadable code and on the other hand slower and more readable code. The problem here is 1. the leaky abstraction of characters and code points and 2. loss of type safety due to C's weak typing rules, which is a bit lost in the "hurr durr messy code."
Do you have a source for this claim?
I find it hard to believe that Microsoft would drop support for ANSI applications.
- Support for weird compilers
- Focus on performance, thus playing "weird" tricks with macros, lookup tables etc.
- especially in headers: protect against users doing crazy sh*t
- standards sometimes change or see extensions
- certainly more :)
First two items can be re-evaluated form time to time, but especially GNU often aims to run "anywhere"
The third item comes from the fact that users can create macros for many things and only some names (leading _ and capital letter or __) are reserved for standards or C++ allowing to overload the comma operator, which can have weird effects. Thus to be conformant to the standards some ugliness is required.
The fourth point combined with different compatibility requirements requires tricks to switch between different implementations based on compiler flags/macros (for instance `#define _GNU_SOURCE` or `#define _POSIX_C_SOURCE=...`)
> an unsigned char or […] EOF
so isalnum is "their own specific variants of character related functions" for uchar, and locale-aware.
This is bullshit. Great, now isalnum returns true for ä if I'm using de_DE.latin1 or whatever, but is still nonsense for any non-western language. And the price we pay for that is a messed up pile of garbage that nobody understands and can segfault.
But all the platforms that use glibc have moved to utf8 over a decade ago. So even for western languages that code doesn't make sense anymore.
I'd argue the usefulness of all that glibc code is very close to zero. Sure, the glibc folks can pat themselves on the back for covering the POSIX spec so well but I really prefer the Linux approach here; follow POSIX where it makes sense but omit the insanity.
Your comparison to Linux makes no sense either. It's not that the code pages themselves are the problem, you can use them for conversion and whatnot. But pretending a collection of functions that is unable to do anything meaningful in the present day and crashes for extra points is fine, because it was written long ago just isn't. The musl solution is correct. If you need to handle strings in a locale-aware manner, use a dedicated lib. Better yet don't use C.
And the crash is a problem with C but it's not a problem with the function. There's a different function, wisalnum, that doesn't crash with out of bounds values, but it also doesn't do Unicode on musl.
The crash is a problem with the function, specifically with glibc. You might even put into question why POSIX defines out of range values as UB instead of requiring it to return false but that's yet another can of worms we better not open here. I fully blame not doing any range checks and doing an oob array access on glibc. UB doesn't force you to create an implementation that crash and burns... Returning false would be great. Maybe abort if you're anal. Have it return something random if you're concerned about speed. But don't freaking crash! The libc should work with the dev, not against them.
It doesn't.
> This is bullshit.
Welcome to POSIX locales.
> Great, now isalnum returns true for ä if I'm using de_DE.latin1 or whatever, but is still nonsense for any non-western language.
Technically it's already nonsense for western languages as ISO-8859 is long outdated, and any codepage other than ISO-8859-1 will not fit inside a uchar when decoded from UTF-8. Also there were non-western languages which fit in 8-bit encodings for which isalnum would work fine.
> And the price we pay for that is a messed up pile of garbage that nobody understands and can segfault.
Welcome to POSIX locales. And also C, where not reading and understanding the implication of every word in the specification means you're bad and therefore deserve everything you get. In this case, POSIX clearly specifies that input values outside of EOF or "unsigned char" is UB.