Fun with Glibc and the Ctype.h Functions
rachelbythebay.com
rachelbythebay.com
> In all cases the argument is an int, the value of which shall be representable as an unsigned char or shall equal the value of the macro EOF. If the argument has any other value, the behavior is undefined.
So this is just a case of glibc being optimized in a way that's really unforgiving if you commit that particular UB.
I'd give up on supporting localization for ctype.
This makes me think, too, "never use ctype, just hardcode my own that assumes ASCII".
https://github.com/mpv-player/mpv/commit/1e70e82baa9193f6f02...
I found it easier to get rid of libc and write freestanding C instead. Linux system calls have none of these problems. This locale bullshit is nowhere to be found. No global state anywhere. No thread-local errno. No stupid stuff like EOF. Writing C became fun again.
IIRC, one nasty area is floating-point numbers, when you're in the locale that doesn't have . as the decimal point, but your program's requirements are such that . is the decimal point.
Like, oh, you're writing a programming language which specifies that.
Like, you know, C does for floating-point constants.
I live in such a country. It's subtle but constant pain even for average users. Even the Windows calculator screws this up: it only accepts either . or , as input instead of both.
> Like, oh, you're writing a programming language which specifies that.
Yeah. The commit message I linked has one such example.
> This is still less bad than that time when libquivi fucked up OpenGL rendering
> calling a libquvi function would load some proxy abstraction library, which in turn loaded a KDE plugin
> which in turn called setlocale() because Qt does this
> made the mpv GLSL shader generation code emit "," instead of "." for numbers
> and of course only for users who had that KDE plugin installed, and lived in a part of the world where "." is not used as decimal separator
Just imagine debbuging this insanity.
It does, however, always accept the . key on the numeric keypad, so many users won't really notice the discrepancy.
“OpenBSD always uses the C locale for these functions, ignoring the global locale, the thread-specific locale, and the locale argument.“
and https://man.openbsd.org/setlocale.3:
“On OpenBSD, the only useful value for the category is LC_CTYPE. It sets the locale used for character encoding, character classification, and case conversion. For compatibility with natural language support in packages(7), all other categories — LC_COLLATE, LC_MESSAGES, LC_MONETARY, LC_NUMERIC, and LC_TIME — can be set and retrieved, too, but their values are ignored by the OpenBSD C library. A category of LC_ALL sets the entire locale generically, which is strongly discouraged for security reasons in portable programs.”
Probably non-standard, but from what I read here, also a defensible choice.
That seems to have disappeared - which supports your view.
Another fun one: With FD_CLR, FD_ISSET, FD_SET you can corrupt memory by merely passing a socket descriptor that is not 0..1024. Pass a negative integer for some undefined behavior as well (shift by negative value occurs here [1])
[1] https://github.com/lattera/glibc/blob/895ef79e04a953cac14938...
This means that if you have a char value (say, an element of a string), you need to cast it to unsigned char before passing it to any of the is*() functions.
glibc needs to solve two hard problems: be very fast and run on innumerable systems. Some of that conditional stuff is because all the world is not Linux or BSD; some of the macrology is there to make sure such handling is performed everywhere needed, and of course the preprocessor is the closest a language like C can get to preprocessing.
I was in the code as glibc started to exist (we paid for a lot of it) and it looked like Musl: very straightforward.
Musl does have a lot less legacy to contend with, and musl is often much slower than glibc, so your point stands, of course.
isalnum works fine of both, it only veers off when you get into UB which is UB.
If you define “works fine” as “gives correct answers even in ub” then musl’s is completely broken since it only gives correct answers for english in ascii.
It can't give correct answers for any non-Latin scripts in any locales.
The problem is ctype and POSIX.
Given that, making ctype only work for ASCII (and maybe EBCDIC if you're really unlucky, which glibc is) is basically sufficient.
Of course, but then why complain that glibc doesn't work outsid of ascii is the point.
That is a misunderstanding on your part. Calling isalnum with out-of-range values is UB. Per-spec: "If the argument has any other value, the behavior is undefined.".
> It is entirely permissible to return 0 for values out of range. Or -1. Or anything else.
Of course it is, it's UB, there's nothing it can't do.
> POSIX is not part of the C language, if a function's behaviour is not defined for a given input, that's not UB
So a behaviour which is not defined is not an undefined behaviour. Sure. Whatever.
> that's a license for the implementation to do whatever is natural. Which can be crashing embarrassingly, but why would you do that.
Why wouldn't you? A table-based implementation is flexible, convenient and efficient, and you don't care what happens outside of the function's bounds.
Is there some other setup I'd need to do to see it work in glibc?
here
#include <ctype.h>
#include <locale.h>
#include <stdio.h>
int main(int argc, char** argv)
{
setlocale(LC_CTYPE, "fr_FR.iso88591");
if(isalpha('ç'))
printf("ok\n");
}
prints ok (with the file in the correct encoding)For example, on my system isalpha(0xe7) is true if I first call setlocale(LC_ALL, "en_US.iso88591").
since when does glibc run on bsd
https://sourceware.org/git/?p=glibc.git;a=blob;f=README;h=b9...
And... yes, glibc does have support for EBCDIC, which is probably ultimately why it has these run-time indirections in its ctype. There's no other reason to have run-time indirections for ctype functions given the limitation of unsigned char values + EOF. That means this code can be simplified a great deal.
Anyways, yes, Drew DeVault's rant misses glibc's need to support EBCDIC, but glibc is exactly like this for every little thing -- an unmaintainable mess. There has to be a better way to produce a fast C library w/o being such a mess on the inside.
Therefore, for instance, isspace(0xA0) might usefully report true if we are in a Unicode locale, otherwise not.
The 0x80-0xFF values are also used in 8 bit extensions over ASCII, like ISO-8859 1 and ISO 8859-15 character sets. E.g. 0xE0 is à in ISO-8859 1 (which is, of course, the same as the Unicode U+00E0 but logically distinct).
A totally different 8 bit extension is KOI-8.
The point is valid that if you don't support any "weird" extensions to ASCII (just ISO Latin) or non-ASCII 8 bit, then there isn't much of a need for run-time table indirection. The cases that may arise can be handled ad hoc. Along the lines of "if we are in an ASCII locale, then report false above 7F, otherwise go through the combined Latin/Unicode combined table".
Instead:
- use UTF-8 locales
- use a Unicode library for all things Unicode
- make your own ctype for when it's Just ASCIIIn a given locale, you can almost certainly regard wchar_t as being a continuation of the range of unsigned char. If you want to know whether the value UCHAR_MAX + 1 is alpha-numeric, you can't pass that to isalnum, but if that value is in the range of wint_t, you can pass it to iswalnum.
For values 0 to UCHAR_MAX, it would be surprising if isalnum and iswalnum produced different results.
Yikes. If you have wide characters, you want iswalnum, or else preprocessing: (ch <= UCHAR_MAX) ? iswhatever(ch) : 0, assuming positive ch.
Illumos: https://github.com/illumos/illumos-gate/blob/9ecd05bdc59e4a1...
...although there is a "sensible" version at:
https://github.com/illumos/illumos-gate/blob/9ecd05bdc59e4a1...
FreeBSD: You have to chase it through "__sbistype" to "__sbmaskrune".
https://github.com/illumos/illumos-gate/blob/master/usr/src/... e.g.