Musings on the C Charter
blog.aaronballman.com
blog.aaronballman.com
However, the C committee refused to define NULL as (void *)0 because they were afraid of breaking existing library implementations that define NULL as 0. They instead introduced a new nullptr keyword that's defined as (void *)0.
I just don't understand the inconsistency.
This was a defect in BSD and the POSIX specification (which literally says it's not intending to contradict the C standard anywhere at the top of every function document), but the committee somehow managed to believe this was cause to deprecate behaviour of realloc() with size zero for 99% of the world. The defect report and associated documents don't seem to be very explicit about the potential impact of the choice to make the behaviour obscelescent in the first place, which might be why this compatibility break was accepted (and eventually completed by making it all undefined).
https://www.open-std.org/jtc1/sc22/wg14/www/docs/summary.htm...
The approach is inconsistent because the committee isn't perfect and every single change to the standard, especially defects, must attempt to best preserve compatibility on a nuanced case-by-case basis, and sometimes when you're in the committee I suspect it's harder to see the wood for the trees. realloc() was broken trying to address a defect report that suggested C's lax definition was potentially a cause of double-frees, a security nightmare, but probably exaggerated in hindsight; and not really made that explicit, although I haven't checked all the minutes and attended the meetings so maybe there was a more conscious decision to deprecate the free-on-zero rule.
> [DR 400] was resolved by loosening the requirements in the standard to allow for the existing range of implementations and included in C17.
> N2438 Clarification Request requested further clarification [...]
> Discussion at the Ithaca meeting and on the mailing list suggested that a call to realloc with a size of 0 be classified as undefined behavior.
> [snip]
> Classifying a call to realloc with a size of 0 as undefined behavior would allow POSIX to define the otherwise undefined behavior however they please.
[0]: https://www.open-std.org/jtc1/sc22/wg14/www/docs/n2464.pdf
There was a discussion if it should be defined as "platform defined" instead, but the wg14 wanted to indicate that they discourage this use of realloc, and undefined behaviour was a way of doing this.
BTW, if the "we" implies you're a WG14 member, I wanted to thank you for helping move the language forward. C23 has tons of nice features that makes it compelling to move to (#embed, true/false being keywords, binary literals, built-in endianness macros, etc), which is why I'm concerned about adding more undefined behavior to the standard because some BSD variants didn't comply with the previous standard.
To be clear: BSD was not breaking the standard, the standard didn't spell out what the correct action was. So different people had different interpretations. The standard states that things that aren't explicitly defined are UB, so one interpretation is that it was UB to begin with.
Is this useful for anything? I wrote my own memory allocator, is there any reason I might want to support this?
NULL is a magic number address that is guaranteed never to be a valid allocation. By allocating an address, you can creat your own magic number, that wont ever be used by any other allocation so you can give it special meaning in your code.
I don't know if changing the realloc() NULL behavior to undefined necessarily breaks code. It definitely leaves the door open to the possibility of breaking code on standard library users but it's still up to the implementation to change the behavior for the code to break.
Looks like gcc does not take advantage of this, but clang does.
This post is a duplicate of "Musings on the C Charter" - https://news.ycombinator.com/item?id=38835994
I notice the following:
- Is currently on the first page while the other one with more points and more comments is not, (at least as I write this)
- Did not get marked as a [dupe], something I see marked within seconds for other posts in similar situations.
- Despite being posted 10 hours after the previous one, instead of linking directly to the previous post, become a separate item.
Can you clarify the rules on this? Is it because of human moderation or other rules?
I agree with most of the author's points but this one goes too far. C is the low-level thing you step out into when you're telling the system "trust me" and taking direct control. Asking people to write raw assembly more often instead of using C is a non-starter.
I understand the underlying issue with "trust the programmer" (it's too parallel with "unsafe by default") but a more nuanced solution than "make it so C is 100% locked down" is going to be required. Even Rust has `unsafe` and pointer-fiddling utilities (but I don't see how a similar idea would make sense for C).
> there should be no invention, without exception
_Generic? The _Atomic mess?
I've been vaguely thinking I need an off ramp from C as it staggers towards a more dangerous and annoying C++ dialect. Currently I think the right play is to write a C99 compiler that doesn't use UB => ruin as an optimising trick that incidentally stomps all over what the programmer thought would happen.
Yes, the post isn't saying that it's gone, it is saying that it should be gone, in the future.
> _Generic? The _Atomic mess?
I do not know the story behind these features, and so your point is a bit confusing to me. That said, I think what you're saying is that these are both inventions, and therefore, this principle isn't true. From the post:
> For starters, we’ve shown we’re perfectly comfortable with invention during the C11 and C23 cycles.
If you have any pointers to the story about standardizing these features, I'd love to hear about them though.
> Currently I think the right play is to write a C99 compiler that doesn't use UB => ruin as an optimising trick that incidentally stomps all over what the programmer thought would happen.
I do think that this is a thing that a lot of people want, but I am unsure that it is a thing that is possible, without seriously inhibiting optimizations. And if you're okay with that, you can compile at -O0. But maybe there is some universe in which you could write different optimizations. Only one way to find out! If you did pull that off, I'm sure your compiler would be popular.
I'm pretty sure whole program optimisation can do the majority of the performance work of the aliasing assumptions. Things like signed integers never overflow are more difficult. My pet theory is that a C compiler that emits machine code that follows a thin layer over the target assembly for semantics will be fast enough for most things C gets used for.
There's been some pretty good quality of life features introduced by recent standards. Just backport them to C99 or something.
the UB part should be limited to the minimal set, now there are too many cases to cause UB.
- strong typed enumerations
- proper arrays with bounds checking (or fat pointers)
- proper strings with bounds checking (could be a library like SDS)
- memory allocation via types instead of sizeof based math
This alone would already fix quite a few problems.
Apple has such Safe C dialect they use for iBoot firmware.
For me the original https link fails with "SSL received a record that exceeded the maximum permissible length. Error code: SSL_ERROR_RX_RECORD_TOO_LONG"
The http link redirected to safebrowse.io/warn.html where there is a "Proceed Anyway" link to get to the content.
I am having trouble parsing that sentence.... so now there's a distinction between the code and the implementation? wat!?
> 1. Existing code is important, existing implementations are not. A large body of C code exists of considerable commercial value. Every attempt has been made to ensure that the bulk of this code will be acceptable to any implementation conforming to the Standard. The Committee did not want to force most programmers to modify their C programs just to have them accepted by a conforming translator.
If you look at the work involved in creating a C89 program that won't compile in modern C, only a tiny fraction of that work involves anything specific to C89. It's almost all about solving some real-world problem in C generally, with the details about C23 compliance being an implementation detail. Aaron seems to view porting a non-compliant C program as loosely equivalent to a security bugfix that changes some low-level details of certain implementations, but doesn't have much impact on the overall code. It might be annoying to fix your K&R declarations / etc, but it won't undo the hard work you did in parsing HTTP headers.
An extreme but illustrative example of this idea: Donald Knuth still maintains TeX in WEB, which is a dialect of Pascal84. He hasn't moved to CWEB or a fancier modern language. I think he's correct: although Pascal84 is fussy and old-fashioned, it is also completely specified, fully stable, and uses a generic ALGOL syntax widely applicable to many modern languages. So the work of translating Pascal84 to modern C can be done with a computer (web2c): TeX's actual WEB source code serves as a complete specification that can easily be translated to a specific implementation. The WEB code is important, but the 1/2/2024 binary output from web2c > gcc is an implementation which is not very important.
> 1. Existing code is important, existing implementations are not. A large body of C code exists of considerable commercial value. Every attempt has been made to ensure that the bulk of this code will be acceptable to any implementation conforming to the Standard. The Committee did not want to force most programmers to modify their C programs just to have them accepted by a conforming translator.
>
> On the other hand, no one implementation was held up as the exemplar by which to define C: It is assumed that all existing implementations must change somewhat to conform to the Standard.
This reads very clearly to me. However,
> Aaron seems to view porting a non-compliant C program as loosely equivalent to a security bugfix that changes some low-level details of certain implementations, but doesn't have much impact on the overall code. It might be annoying to fix your K&R declarations / etc, but it won't undo the hard work you did in parsing HTTP headers.
I do think you're getting at something here, which is the change Aaron is advocating for. I think it's a bit more subtle than that though: most people think of C as being 100% backwards compatible. But that's not really true. Famously, for example, C++ removed the gets function outright. I think Aaron isn't arguing that backwards compatibility should be thrown out entirely, but that the committee should build in a formal process for deprecation and eventual removal, and that that is better than the status quo. And that this power should be used wisely, not making changes that would be a burden on users for little gain. Just that it is okay to give users a little burden for big gains, rather than the current status of "basically no burden at all, but also improvements are difficult."
So I was summarizing Aaron's gist, and in particular his comments about wanting to avoid a Python 2/3 schism. This was a case of "big burden for big gains," the burden being so large that developers didn't want to harm their hard-earned business logic implementation because 2to3 missed an edge case in a dict comparison.