My review of the C standard library in practice
nullprogram.com
nullprogram.com
char buf[N] = {0};
fread(buf, N-1, 1, f);
puts(buf);
This doesn't work reliably because the arguments are wrong. It's "fread(buf, size_of_item, number_of_items, file_ptr)". You're telling it to read one item of size N-1. If there's not enough data to read for a complete item, the rest of the "item" might be written with garbage. If you do it the right way around - reading N-1 items of size 1, you actually get the result you expect.Is this an unnecessary footgun in the standard C library API? Probably yes. But if you complain about it, I really rather you complain about the actual problem.
Other than this, there's a few things I would agree with, and then a few other things that are issues with _specific_ standard C libraries, e.g. assert just exiting the process. Mixing C library API / POSIX problems with specific C library problems isn't particularly great for an article like this.
Anyway, the article registers appropriately "uncooked" for someone reinventing pkg-config; as far as my personal "reputation database" is concerned this is strike 2 for the author.
https://man7.org/linux/man-pages/man2/read.2.html
ssize_t read(int fd, void *buf, size_t count);
That's how it works in every sane I/O API I've seen. What's even the point of the size_of_item argument?However, this functions as a red flag for the article itself, as in: in an article dismissing something with very little nuance, finding one of the arguments exhibit a lack of good understanding of that thing raises an interrupt for me that its other arguments may also be flawed.
The article overall just rubs me as hubris. If you go about dismissing decades of engineering history, you better get your facts right. It's easy to write things like this when you're at the wrong end of the Dunning-Kruger bathtub, so you have to show you are in fact at the right end of it. (While in the middle, people tend to not write articles like this in my experience.)
It seems so simple. Yet certainty eludes me.
"Dunning-Kruger bathtub"
This phrasing is gold btw.
(Not only did Dunning-Kruger not even claim that incompetent people believed they were confident, but the entire study doesn't reproduce.)
"Many criticisms of the Dunning–Kruger effect have the metacognitive account as their main focus but agree with the empirical findings themselves."
https://en.wikipedia.org/wiki/Dunning%E2%80%93Kruger_effect#...
Would it even matter? It's regularly used to describe a common phenomenon. You understood what the Parent Comment was describing, as did I. For that reason alone, it's effective communication.
Wow, so cocky.
> If you go about dismissing decades of engineering history, you better get your facts right.
I find most of the opinions on that page are well explained, and they aren't spectacularly bold anyway. You will be hard pressed finding people that think locales are usable and most of string.h is not historical baggage.
I also learned something new, for example the thing about isXXX() taking unsigned values.
> finding one of the arguments exhibit a lack of good understanding of that thing raises an interrupt for me that its other arguments may also be flawed.
You could try the thing about reading in a benevolent way. Don't choose the interpretation that would most upset you.
From what I can see you found one nit where the author might not have a complete understanding on the issue, or maybe they just swapped the two arguments. (I'm not quite sure what that explaination was about, and I didn't bother to research. You're right of course about how fread should be used). And then you went back to HN to tear the article in pieces (btw. I find the post to be well above average quality).
As far as I'm concerned, this is strike 1 on my list.
> You could try the thing about reading in a benevolent way. Don't choose the interpretation that would most upset you.
Et tu, Brute?
Probably, yes. In the earliest versions of C, stdio hadn’t been invented yet, and all IO was done using Unix file descriptors - which was fine, because C only ran on Unix. But, they were keen to prove the language was cross-platform, and Bell Labs used both IBM and Honeywell mainframes, so they started porting C to those two platforms, as a proof-of-concept of its portability. And there they ran into a big problem-mainframe IO was radically different from Unix, and the file descriptor API was rather ill-suited to it. So, Mike Lesk invented a “portable IO package”, which abstracted over the differences between mainframe and Unix IO-and that package was the ancestor of stdio.
And I think that explains why fread() is designed the way it is. IBM mainframe IO is fundamentally record-oriented [0] (and probably Honeywell was too, although I know far less about it). To Unix, reading 10 records of 80 bytes each, 80 records of 10 bytes each, or 1 record of 800 bytes each - the three are largely equivalent. Not so under MVS-if, when you create a file, you declare it as being composed of 80 byte fixed length records, the OS will force all reads/writes to be multiples of 80 bytes-hence why fread() makes you declare the record size, since it is much more essential information on that platform.
[0] Historically speaking; nowadays z/OS, the successor to MVS, supports Unix-style byte-oriented IO as well as classic mainframe record-oriented IO; but back in the 70s, those Unix compatibility features were still 20 years away. And classic mainframe apps still predominantly rely on record-based IO
Amusingly, only the most simplistic graphs of the Dunning-Kruger efect would portray it as matching a bathtub curve. In particular the rather slow slope of confidence rising as competence improves is dramatically (and significantly) different from the sharp drop-off of false confidence. Rather than two steep sides, there is a marked asymmetry.
I suspect that you may have misunderstood the author's point though. Even if you swap the arguments as you suggest, you still run into this problem described in the link that the author provides:
When used on a text mode stream, if the amount of data requested (that is, size \* count)
is greater than or equal to the internal FILE \* buffer size (by default the size is 4096 bytes,
configurable by using setvbuf), stream data is copied directly into the user-provided buffer,
and newline conversion is done in that buffer. Since the converted data may be shorter than the
stream data copied into the buffer, data past buffer[return_value \* size]
(where return_value is the return value from fread) may contain unconverted data from the file.
For this reason, we recommend you null-terminate character data at buffer[return_value \* size]
if the intent of the buffer is to act as a C-style string.
i.e. his complaint is that if he initialises the buffer to zeros and reads N bytes into it there is the possibility that zeros after the N bytes are overwritten.This is distinct from the problem of not getting enough information back if you swap the argument order, and as it matches his decription in the text before the link I would assume that the swapped arguments are actually a typo rather than a misunderstanding of the API.
items = fread(array,sizeof(array[0]),count,fp);
vs: items = read(fh,array,sizeof(array[0]) * count);
That multiple could overflow. fread() (and fwrite()) return the number of items read, not the number of bytes (or characters). The standard also points out that if size is 0, or nmemb is 0, nothing is done (and 0 is returned).No it isn't. What C standard library are you using? Above code is correct in my C standard library ... man fread:
SYNOPSIS
#include <stdio.h>
size_t fread(void *ptr, size_t size, size_t nmemb, FILE *stream);
....
The function fread() reads nmemb items of data, each size bytes long, from the stream pointed to by stream, storing them at the location given by ptr."The fread() function shall read into the array pointed to by ptr up to nitems elements whose size is specified by size in bytes, from the stream pointed to by stream. For each object, size calls shall be made to the fgetc() function and the results stored, in the order read, in an array of unsigned char exactly overlaying the object. […] If a partial element is read, its value is unspecified."
It doesn't say you stop when fgetc() returns EOF (or an error), so by a strict reading the remainder of partial items is filled with fgetc()'s "EOF" return value. But the more relevant thing is that it says the value of a partial element is unspecified.
(Edited to point out the "unspecified" aspect.)
(Also, check the return value...)
The objection is to overwriting N-1 bytes when only N-2 bytes were read.
Which... doesn't sound completely unreasonable?
On the other hand (especially with SIMD CPUs, but maybe also with simpler instruction sets) I can understand that copying blocks of e.g. 8 bytes at a time might be faster than the special case at the end of a read for copying only 7 bytes. (Because you need to follow an "unlikely" branch to get to the 7-byte code which stalls the CPU, maybe.) So if the output buffer has space for it, it might be more performant to copy 8 bytes including a junk byte and just say you copied 7, than copying only 7 bytes.
And the author does seem to object to a number of other design decisions which could hurt performance.
Completely agree: freestanding C is a superior language. I'm so happy to discover I'm not alone in thinking like this. The author is very thorough in his criticism of libc, I learned a lot from the post.
> The platform code is small in comparison: mostly unportable code, perhaps raw system calls, graphics functions, or even assembly.
For me this is what made programming fun again! I don't even bother with multiplatform implementations anymore, I go straight for Linux system calls. Turns out to be a much better interface compared to libc.
> On some platforms it will still link libc anyway because it’s got useful platform-specific features, or because it’s mandatory.
If anyone would like to know why such a thing would be mandatory, I began writing about this subject literally just a few days ago.
https://www.matheusmoreira.com/linux/system-calls
Still very much a work in progress but it does provide context for his claim.
It's sort of amazing how little userspace code a crude webserver using Linux syscalls can be: https://github.com/jcalvinowens/asmhttpd
One 4K page of code! No stack!
How much traffic can it handle, have you benchmark it?
FROM alpine as build
RUN apk add --no-cache build-base gcompat nasm
RUN mkdir /src
COPY \* /src
RUN cd /src \
&& make
FROM scratch
COPY --from=build /src/asmhttpd /asmhttpd
CMD ["/asmhttpd", "/http"]
(Run: `podman build -t $IMAGE .; podman run -p 8080:80 -v $PATH:/http -d $IMAGE`)I agree with many points from the article but strongly disagree with this one.
It makes much more sense to treat a null pointer as an invalid, not present, object. Treating null pointer as an empty string (perhaps by expecting them first N pages to contain just \0 bytes, as various older Unix systems predating Linux did) feels like a recipe to hide bugs.
For example, I want to distinguish the output from getenv of "the variable was found but has an empty value" from "the variable was not defined".
If you have some random struct Foo, it's very clear that a null Foo pointer just means "no valid Foo value" and attempting to dereference it should crash (rather than optimistically return ~random data and naively continuing program execution, making bugs harder to detect).
I expect any reasonable programmer would suggest treating a string (or array of any types) identically.
Also it's "its", not "it's".
Empty string is not a 'zero-sized object'. It's an object with size of one byte that is zero.
There is nothing worong in treating [NULL, NULL) as a valid, 0-sized range. As a sibling comment pointed out, this is already valid in C++ for std::string_view. Or more appropriately, for std::copy as well.
TFA says that is the case for some libc functions. I agree that this is unreasonable. Eg. memcpy shouldn't care if pointers it should not follow are invalid; a copy of length 0 ought to be spec'd as a no-op.
Many of these cases were due to backward compatibility - different pre-existing implementations did different things, everybody wanted to claim conformance to the standard but nobody wanted to change their existing behaviour to do so (potentially breaking their existing application base), so the political compromise was to make them UB.
Thankfully, the ISO C standards committee, in recent years, seems less afraid to make breaking changes than they used to be. And many of these UB cases were to placate implementations that are now long-dead-sometimes you’ll find that despite being UB, most or all major contemporary implementations do the same thing, which increases the odds that a future release of the standard might standardise those behaviours. But this is what happens when a programming language gets to be 50 years old.
To treat a C string as empty, I have to dereference the pointer and see that the first byte is \0. If the pointer is 0x0 then I crash.
But if I want to copy 0 bytes to or from 0x0, I do not have to dereference 0x0.
I would want to spend way too long with the standard before I opined on the correctness of the author's statement, but "zero-sized object" is definitely different from "zero-length string".
The libc is the weakest point of the C programming language, outdated, bad naming conventions, and in some cases harmful APIs.
It also has its merits, simplicity and availability, it is good as a fallback, but I think it is generally preferable to have a full API coverage when using a framework, including basic things such as printf/malloc/fopen as there are better ways to make those and it is quite important to maintain similar convention across the whole codebase.
What's the rationale of providing this as a random set of code snippets instead of putting together a library?
For sure, there's no legal expectation of support. If you don't support your open-source project, I can't sue you.
But there is a social expectation of support. If the issues page is a ghost town, proposed patches just sit, and there hasn't been a release in years, we call the project "unmaintained" and discourage its use. At that point, if there's interest, somebody can fork it and support the fork, but forking is a serious act: it strongly signals that you don't like the direction that the current maintainer is heading. If that project is truly unmaintained, that's fine, but all-too-many open-source projects end up with rival forks when the original maintainer comes back.
So yeah. There's zero legal expectation of support. But the social expectations are real, and providing a snippet (at least in my mind) has less of an expectation than a library.
[1] https://lists.ozlabs.org/pipermail/ccan/2022-September/00141...
A lot of effort seems to go into stopping people from using some of the less well designed parts.
I also mostly try to avoid the C standard library because I can’t be bothered to remember which functions are flawed and which are (mostly) safe to use. But I really don’t like to write my own platform-specific wrappers for e.g. converting between string encodings.
If you want your compiler to know what memcpy does, so that, for example, it can compile it as an inlined “load register, store register” instruction pair _if_ it knows it’s copying 8 bytes, you have to #include<string.h> (with the angle brackets; #include "string.h" won’t do)
So, if you want optimal performance, you can’t do without the compiler header.
You’re still free to link with your own implementation of memcpy, of course, but don’t expect it to be called in all places where the source code contains calls to it.
But just like most pitfalls in the language, it's a consequence of the early developers not putting much thought when assigning "namespaces" on the C library functions/macros.
As for the locales non-standard behavior in GNU tools, it's no strange that GNU project always historically deviated from POSIX standard.
Of course you want it in there in case you need it, but writing so much about them when you say "standard library in practice" is a bit suspect.
The C stdlib is essentially an SDK for very simple 70's UNIX-style command line tools, but operating systems have moved on, while the C standard library is unfortunately stuck in the past.
(IMHO a "C stdlib v2" would be much more important than any new language features, but it doesn't look like the C committee wants to go down that road).
musl and glibc are implementations of libc, not alternatives to it.
All you need are the system calls of your kernel. You don't need libc to call those.
As a small example: `strtok` is simultaneously a very useful function and one that's annoying to use correctly (much less independently implement correctly). It isn't a system call.
As for scanf: I want it to either succeed in parsing all of the arguments, or fail, but in this case leave the file position unchanged (so random access is required on input). This would allow trying to parse with another call to scanf, but with a different format string. For lists involving an arbitrary number of items, provide a conversion specifier that leaves the input alone and returns success: this way you can iterate the list.
For example, these would be equivalent:
if (scanf(f, "%d %d %d", &a, &b, &c)) printf("Success!\n");
if (scanf(f, "%d %e", &a) && scanf(f, "%d %e", &b) && scanf(f, "%d", &c)) printf("Success!\n");
/* %e means more input expected, without it scanf fails if no EOF at that point */
I should be able to provide my own parsing functions in it: imagine a conversion specifier for JSON that returns a tree.The existence of sprintf and sscanf point to bad design: I should be able to open a block of memory as a FILE, which would make sprintf and sscanf redundant. Also I should be able to open some kind of block device: meaning the FILE has pointers to user provided read and write functions. With this I could printf and scanf to some indirect device such as an EEPROM.
We should be using scanf for the program's argument list. I mean the program should not get an array of pointers, but instead a FILE that contains the arguments. The memory layout is then abstracted. The pre-conversion of the command line to an array is totally unnecessary, only exists because libc's scanf sucks.
Arbitrary length strings could have been handled with FILEs..
> As for formatted input, don’t ever bother with scanf.
With an entire other article about scanf as the link.
edit:
As for your other point, files may not even be seekable, let alone random access. How should that be implemented for a pipe?
If the input is really not random access capable, provide buffering. You could imagine a version of FILE that automatically does this. There could be a some kind of truncate call to make it forget all past input, once you are sure it's not needed.
In an OS you will have the resources to do this (malloc..). In other contexts (embedded), you may not.. but likely in those cases you don't have pipes.
With a bit more we can have full arbitrary backtracking parsing. This parses "<hex> bob | <decimal> fred" or prints the unknown input:
if(
fscan_push(f) && fscanf(f, "%d %e", &n) && fscanf(f, "fred") && fscan_done(f) || fscan_backup(f)) printf("got fred(%d)\n", n);
else if(
fscan_push(f) && fscanf(f, "%x bob", &n) && fscan_done(f) || fscan_backup(f)) printf("got bob(%x)\n", n);
else
printf("got something else '%f'\n", f);
'%e' expect more inputfscan_push: save current location on a stack associated with FILE. Return true.
fscan_done: drop location from stack. Return true. If stack becomes empty, drop input history up to current file position.
fscan_backup: restore to location from stack. Return false.
'%f' in printf prints the rest of the specified FILE on the stdout.
C23 fixes this - it introduces %w32d as a replacement for PRId32, %w64d to replace PRId64, %w64u, etc.
A good qsort or bsearch mechanism would allow the comparisons to be inlined. Performance critical code can't use the libc versions due to this.
But yes, I have seen numbers comparing qsort and std::sort. There is no competition in this. The latter is pretty much always vastly better in wall clock time due to inlining.
Yes you trade code size for that, but today memory and storage is cheap.
IMO libc's should offer an extern inline definition of those functions, so the compilers are able to inline or function specialize them when needed.
I remember the days when an optimizer didn't really optimize a function pointer call. Now they can probably pull it off if it's the same compilation unit. And if you have link time optimizations, I've seen it do inlines across object files, but obviously that won't work for dynamically linked libc.
But a static inline qsort in a libc header is probably a good call.
But what is the alternative? Rolling your own equivalents?
If you need cross-platform, depending on your application it may be reasonable to create wrappers for your string usage with platform-specific implementations that use the best option available on each OS and fall back to POSIX.
But then again, I'm not a C dev by trade.
setjmp and longjmp will have to stay, though.
I practically almost never use those.
For me, I have a stack allocator that knows about cleaning up, as well as about setjmp and longjmp, so I have a function I can call that will deallocate everything in the stack allocator and do the jump. It's basically a safe exception because everything will be cleaned up properly (if I allocate everything on the stack allocator or be rooted in the stack allocator, which I do).
Links to it are in my comments, but I wouldn't look at it.
Instead, I'll describe something you can implement with helps that would be much simpler.
Keep a stack of data about allocations. You should keep a pointer to the allocation, the size, and the "destructor" (a function pointer to a function that takes a void pointer and returns nothing).
Write 3 destructors. Two should do nothing, but they should be separate. One should be for scopes, and the other should be for functions. The last destructor should cast its void pointer to a jmp_buf pointer and do a longjmp() on it.
Then have a normal allocation and deallocation function on the allocator. This is for normal allocations.
Have a function to enter a scope, enter a function, and set a jump. In each case, they should push data onto the stack with the appropriate destructor. (I push a pointer to the function name for functions, for debugging.)
Then have functions to exit a scope, exit a function, and do a jump. These should "unwind" all allocations until they reach an allocation that uses the scope destructor, function destructor, and setjmp() destructor, respectively. (You can text if two function pointers are equal to each other.)
I'll be happy to answer any more questions.
/jk
1. There is no "industry standard" for C compilers, for instance in the game dev world, the closest thing to an "industry standard compiler" is MSVC, and not GCC or Clang.
2. While Clang tries to emulate some GCC-isms, it also has important differences to GCC. On Windows, Clang also emulates a lot of MSVC-isms, so emulating GCC really doesn't mean anything.*
https://www.godbolt.org/z/Ezb7hh6eq
...and for the same reason, AFAIK the compiler isn't required to preserve struct padding bytes either in such a case, so even though memcpy() is used for copying one struct into another in the original source code, the compiler may actually generate code which behaves differently from a straight memcpy() - but in this case I'm not 100% sure, I just wouldn't depend on it, because I've been bitten in the ass multiple times when it comes to assuming the content of padding bytes.
He mentions the problems caused by the criminal use of globals in locales. That problem has lead to numerous blog posts that have hit the HN front page.
This one comes to mind: https://aras-p.info/blog/2022/02/25/Curious-lack-of-sprintf-...
What does it mean?