C23 Implications for C Libraries
htmlpreview.github.io
htmlpreview.github.io
Who exactly are these new C-standards for?
I interact and use C on an almost daily basis. But almost always ANSI C (and sometimes C99). This is because every platform, architecture, etc has at least an ANSI C compiler in common so it serves as the least common-denominator to make platform-independent code. As such it also serves as a good target for DSLs as a sort of portable-assembly. But when you don't need that, what's the motivation to use C then? If your team is up-to-date enough to quickly adopt C23, then why not just use Rust or (heaven forbid, C++23)?
I'd love to hear from someone who does actively use "modern" C. I would love to be a "modern C" developer - I just don't and can't see its purpose.
An example: The C11 memory model + <stdatomic.h> + many compilers supporting C11 has/had a positive impact on language runtimes. Portable CAS!
> If your team is up-to-date enough to quickly adopt C23, then why not just use Rust or (heaven forbid, C++23)?
Another example: If you're programming e.g. non-internet-connected atomic clocks with weather sensors like those produced by La Crosse, then there's no real security model to define, so retraining an entire team to use Rust wouldn't make much sense. (And, yes, I know that Rust brings with it more than just memory safety, but the semantic overhead comes at a cost.)
Another example: Writing the firmware to drive an ADC and broker communication with an OS driver.
Another example: The next Furby!
[1] https://en.wikipedia.org/wiki/C11_(C_standard_revision)#Opti...
I prefer the builtins. With _Atomic you can easily get seq-cst behavior by mistake, and the <stdatomic.h> interfaces are strictly speaking only valid for _Atomic types.
Portability isn't binary. It's the result of work being done behind-the-scenes to provide support for a common construct on a variety of hardware and operating systems. It's a spectrum.
GCC certainly is portable, but it doesn't support every ISA and OS. Over time it has even dropped support for several.
Random thoughts since I'm still in the process of waking up…
1. Most of <stdint.h> is optional
2. long is 64 bits on Tru64 Unix which is valid under all versions of the standard
I can easily whip up CAS for all platforms I care about, plus a few more for bonus, using nothing but C99 or C90 (and have actually done it before).
Possibly, without even using any language extensions, unless I want the operations inlined.
I'd be surprised if the former were possible. From a quick web search, C doesn't even give you the guarantees necessary for Peterson's Algorithm, [0][1] and volatile doesn't help. [2][3]
[0] https://codereview.stackexchange.com/a/124683
[1] https://stackoverflow.com/questions/35527557/errors-with-pet...
[2] https://web.archive.org/web/20160525000152/https://software....
[3] https://old.reddit.com/r/programming/comments/d457c/volatile...
I guess if I could rephrase my original question: People who are going to adopt C23 - who are you and what field/industry/line-of-work are you in?
Asking because in my line-of-work, C is ubiquitous and I personally love coding in C, but anything beyond ANSI C (or C99) is "cool" but undermines the point of C as I've used it, which is its use cross-platform for a huge set of common and uncommon architectures and instruction sets. If something only needs to run on common, conventional platforms, C, however much I love it, would no longer necessarily be a strong contender in light of many alternatives. It seems like these standards target an ever shrinking audience (much smaller than the whole universe of software developers working in C).
> ANSI C is the de-facto baseline everything has in common,
> using C11+ narrows which platforms you're able to build for.
Sure, but which platforms are you thinking of? I think you're
overestimating the proportion of C codebases that care about targeting
highly exotic hardware platforms. GCC can target all
sorts of ISAs, including embedded platforms.
It can't target, say, the 6502, or PIC, or Z80, but they're small niches.I'm not an embedded software developer though, perhaps there are more developers using niche C compilers than I realise.
> Even if C11 does provide the primitives for correct implementation
> of Peterson's Algorithm, what about the other platforms that don't
> have a C11 compiler - it does them no good.
If portable atomics/synchronisation machinery is not offered by your
C compiler or its libraries, I figure your options are:1. Use platform-specific atomics/synchronisation functionality
2. Leverage compiler-specific guarantees to write platform-specific atomics/synchronisation functionality, if your compiler offers such guarantees
3. Write your atomics/synchronisation functionality in assembly, and call it from C
Here's a project that uses all 3 approaches. [0] (It's old-school C++ rather than C, but it's the same in principle.)
I'm fairly sure it's not possible to implement your own portable synchronisation primitives in standard-compliant C90 code. As I understand it, the C90 standard has nothing at all to say on concurrency. It's possible that such an attempt might happen to work, on some given platform, but it would be 'correct by coincidence', rather than by definition. (Again, unless the particular compiler guaranteed the necessary properties.)
[0] https://github.com/gogglesguy/fox/blob/fe99324/lib/FXAtomic....
You probably wouldn't want to build an atomic compare-exchange in standard C, even if it were possible; you find out what the hardware provides and work with that.
> The code wasn't standard C, and it didn't just "happen" to work
Thanks, that makes sense.At the risk of sounding pedantic, you did say using nothing but C99 or C90, implying use of standard features only.
> You probably wouldn't want to build an atomic compare-exchange in
> standard C, even if it were possible; you find out what the
> hardware provides and work with that.
Agreed. bool cmpswap64(uint64_t *location, uint64_t if_old, uint64_t then_new);
Code that relies on the header doesn't have to process anything compiler-specific. FFI could be used to bind to that function from non-C languages.IMHO, both C++ and Rust feel too much like puzzle solving ("how do I solve this problem in *C++*" or "how do I solve this problem in *Rust*?"), when writing C code, the programming language disappears and it simply becomes "how do I solve this problem?").
PS: I agree that the C standard isn't all that relevant in practice though, you still need to build and test your code across the relevant compilers.
In C, I feel like I'm building a house out of tinker toys, C++ is Lego Techniks, and Rust I'm using bricks and mortar. FWIW, Python feels like waterballoons and drywall to me; while it might look OK from the outside, one thing pierces your exterior and things tend to sag sadly from there.
Don't use (most of) the C stdlib, it's useless and hopelessly antiquated (not just for string manipulation), instead use 3rd party libs.
> who owns what is something I still have to decide except that the compiler actually helps out
Lifetime management for dynamically allocated memory in C should also be wrapped in libraries and not left to the library user. In general, well designed C library APIs replace "high level abstractions" in other languages.
But I agree, the memory safety aspect of Rust is great, and I'd love a small language similar to C, but with a Rust-like borrow checker (but without all the stdlib bells'n'whistles that come with it, like Box<>, RefCell<>, Rc<>, Arc<>, etc etc etc...) - instead such a language should try to discourage the user from complex dynamic memory management scenarios in the first place.
It's not the memory safety in Rust that turns me off, but the rest of the 'kitchen sink' (for instance the heavy functional-programming influence).
And now I have more problems ;) . (And I develop on CMake; consistently using external deps is a nightmare.)
The thing is that I usually am that library author (or working in an area that acts like that).
I'm not sure what you expect to be left if you say "the stack is all you can use" (which is what I understand to be remaining when you remove those "bells'n'whistles").
I also really enjoy the functional aspects. I don't want to think about what some of the iterator-based code I've done looks like in C or C++ (even with ranges).
Not what I meant, heap allocations are allowed (although the stack should be preferred if possible), but ideally only with long lifetimes and stable locations (e.g. pre-allocated at program startup, and alive until the program shuts down), and you need a robust solution for spatial and temporal memory safety, like generation-counted index handles: https://floooh.github.io/2018/06/17/handles-vs-pointers.html).
This statement very much resonates with me. It's honestly one of the things I like about C. Although it's not perfectly like this for me all the time. For example string manipulation is not great.
An other aspect I like about C is there is not a plethora ways of doing the same thing which I have found always made it more readable than rust and C++.
I do dial up the warnings to 11, yet it is not enough.
I've written C code that's currently running on hardware orbiting the earth. I'll never do it again if I can help it; it wasn't worth the stress. You only get one chance to get it right.
I guess in such a situation I would not not trust any compiler (for any programming language), no matter how 'safe' it claims to be, but instead carefully audit and test every single assembly instruction in the compiler output ;)
You cannot rely on that. If you're maintaining C90 code, with a C90 compiler or compilation mode, you should go by what is written ISO 9899:1990, plus whatever the platform itself documents.
We actually don't want compiler writers mucking with the support for older dialects to try to modernize it. It's a backward-compatibility feature; if you muck with backward compatibility features, you risk breaking ... backward compatibility!
Also, most of these corner cases are so obscure that the vast majority of people with decades of C experience have not encountered them. C is an extremely explored space.
As a random example for something as fundamental as classes in C++, the page https://en.cppreference.com/w/cpp/language/classes shows ten defect reports.
As for new developments, see for example https://news.ycombinator.com/item?id=33675462 which uses C11.
For me and for many colleagues in my lab? C is quite big in scientific computing and signal processing. Fortran would be slightly better, and it is widely used, but not directly around me. The C99 standard, which added complex numbers and variable length arrays, was truly a godsend in the field. I cannot imagine working without it.
If you write a numerical algorithm that needs to be run 15 years from now, then C and Fortran are possibly the sanest choices. If you do something in other, fancier, languages, you can be sure that your code will stop working in a few years.
The new C standards are really minor changes to the language, and they happen in the span of a decade. It is quite easy to be up to date. And in the rare case that your old code stops compiling, the previous (less than a handful) versions of the language are all readily available as compiler options in all compilers. You can be reasonably sure that a C program written today will still compile and run in 20 years. You can be 100% sure that a python+numpy program won't. If you care about this (for example, if you are writing a new linear algebra algorithm to factor matrices), then choosing C is a rational, natural choice.
It's possible to use a phyton+numpy program in 20 years, but you also have to save the entire environment and make sure that it works air-gapped (otherwise external dependencies would fail). One possiblity would be to store it as a qemu virtual machine. It's very possible today to boot up stuff as VMs that is 20 years and older (e.g. 20 year old Linux distros or Windows XP iso from early 2000s).
There's a lot of reasons to use C23 over rust
- multiple compiler implementations
- works on more platforms
- defined standard
- ability to create self-referential data structures without hacky workarounds
- immediate, easy access to large numbers of C libraries
(For the record I like rust, but the evangelism over the past half decade has been pretty ridiculous. Consider this counter propaganda).
> Combined with verification costs that don't vary that much from
> language to language, and there's a long future where, for
> safety-critical applications, there's no downside to C -- the cost
> of verification and analysis swamps the cost of writing the code
That doesn't sound right. You really want to get the code right early on. The later bugs are discovered, the more costly the fix. You may have to restart your testing, for instance.If the language helps you avoid writing bugs in the first place, that should translate to quicker delivery and lower costs, as well as a reduced probability of bugs making it to production. The Ada folks are understandably keen to emphasise this in their promotional material.
> the cost of qualifying a new language's toolchain would be absurd
As I understand it, this typically falls to the compiler vendor, not to the people who use the compiler. A compiler vendor targeting safety-critical applications will want to get their compiler certified, e.g. [0]. To my knowledge we're nowhere near a certified Rust compiler, although it seems some folks are trying. [1]To use my baby as an example: free_sized(void *ptr, size_t alloc_size) is new in C23. I can detect whether or not it's available and use it if so. If it's not available, I can just fall back to free() and get the same semantics, at some performance or safety cost.
We emailed about this a little contemporaneously (thread "Sized deallocation for C", from February), and I think we came to the conclusion that glibc can make interposition work seamlessly even for interposed allocators lacking free_sized, by checking (in glibc's free_sized) if the glibc malloc/calloc/realloc has been called, and redirecting to free if it hasn't. (The poorly-named "Application Life-Cycle" section of the paper).
Spec says it's functionally equivalent to free(ptr) or undefined:
If ptr is a null pointer or the result obtained from a call to malloc, realloc, or calloc, where size size is equal to the requested allocation size, this function is equivalent to free(ptr). Otherwise, the behavior is undefined
Even the recommended practice does not really clarify things:
Implementations may provide extensions to query the usable size of an allocation, or to determine the usable size of the allocation that would result if a request for some other size were to succeed. Such implementations should allow passing the resulting usable size as the size parameter, and provide functionality equivalent to free in such cases
When would someone use this instead of simply free(ptr) ?
It's a performance optimization. Allocator implementations spend quite a lot of time in free() matching the provided pointer to the correct size bucket (as to why they don't have something like a ptr->bucket hash table, IDK, maybe it would consume too much memory overhead particularly for small allocations?). With free_sized() this step can be jumped over.
There are four big reasons why:
* Atomics. These are the biggest missing feature in older C.
* Static asserts. I can't tell you how much I love being able to put in a static assert to ensure that my code doesn't compile if I forget to update things. For example, I'll often have static constant arrays tied to the values in an enum. If I update the enum, I want my code to refuse to compile until I update the array. I have 20 instances of static asserts in my current project.
* `max_align_t`. It's super useful to have a type that has the maximum alignment possible on the architecture.
* `alignof()` and friends. It's super useful to get the alignment of various types. Combined with `max_align_t`, it is actually possible to safely write allocators in C. Previously, it wasn't really possible to do safely or portably. And I have at least three allocators in my current project.
You're right that C11 doesn't have nearly the reach the ANSI C does, but it does have slightly more than Rust, much more if you consider Rust's tier 3 support to be iffy, which I do.
And it does have one HUGE advantage against Rust: compile times. On my 16-core machine, I can do a full rebuild in 2.5 seconds. If I changed one file in Rust, it might take that long just to compile that one file.
That's not to say Rust is without advantages; one of my allocators is designed to give me as much of Rust's borrow checker as possible, on top of API's designed around that fact.
tl;dr: I use modern C for a few features not found in C89, for the slightly better platform support against Rust, and for the fast compiles.
And I certainly appreciated C11 when writing Objective-C, so I’m sure people with large codebases of ObjC will appreciate it (though most will be using Swift for new features nowadays).
Example:
https://pasteboard.co/VkjrJIOZzaiR.jpg
(Sorry for pasting code as an image, I'm on my phone)
When you have a C code base or experience with C those features may be enough not to make a complex transition.
Having a simple tool evolve a bit may be what you need as opposed to making the change to a much more complex tool.
This text is still in force, it seems:
“ 13. Unlike for C99, the consensus at the London meeting was that there should be no invention, without exception. Only those features that have a history and are in common use by a commercial implementation should be considered. Also there must be care to standardize these features in a way that would make the Standard and the commercial implementation compatible. ”
I read this as saying that anything that gets standardized should be available in one of the major implementations. In practice, most of the qualifying features will have been implemented in both GCC and Clang in the same way, so for most users, there is not much benefit from standardization. Some may feel compelled to support the ”standard” way and the “GCC/Clang” way in the same sources, using a macro, but that isn't much of a win in most cases. Of course, there will be shops that say, “we can't use feature until it's in the standard”, but that never really made sense to me.
Things are considerably murky on the library side. In my experience, library features rarely get standardized in the same way they are already deployed: names change, types change, behavioral requirements are subtly different. (Maybe this is my bias from the library side because I see more such issues.) For programmers, the problem of course is that typical applications do not get exposed to different compiler versions at run time, but it's common for this to happen with the system libraries. This means that the renaming churn that appears to be inherent to standardization causes real problems.
Others have said that new standards are an opportunity to clarify old and ambiguous wording, but in many cases the ambiguity hides unresolved conflict (read: different behavior in existing implementations) in the standardization committee. It's really hard to revise the wording without making things worse, see realloc.
So I'm also not sure what value standardization brings to users of GCC and Clang. Maybe it's different for those who use other compilers. But if standardization is the only way these other vendors implement widely used GCC and Clang extensions (or interfaces common to the major non-embedded C libraries), then the development & support mode for these other implementations does not seem quite right.
For the C version you have __STDC_VERSION__, but there's no similar facility to check if e.g. J.5.7 is supported, which effectively makes the behavior that's explicitly omitted in 7.22.1.4 and 6.3.2.3 go from "undefined" to supported by C23 + the extension.
I understand why C can't have some generic "is this undefined?" test, but it seems weird not to be able to ask if extensions defined in the standard itself are in effect, as they define certain otherwise undefined behavior. The effect is that anyone using these extensions must be intimately familiar with all the compilers they're targeting.
This pleases me greatly. Two's complement won decades ago. This also means they could define integer overflow as 2's complement rollover, which is almost universal but is still considered undefined behavior.
For example almost never does a programmer want int to rollover in a for loop. Defining that behavior doesn't help, and indeed makes tools like linters or sanitizers less useful.
While two's complement signed rollover can be pretty useful, in 99% of cases you don't need it, and casting to unsigned for the 1% is a worthwhile sacrifice for the better optimizations for the common case, and it'll be pretty clear when you want wrapping or not (and sanitizers will be able to keep reporting signed overflow as errors without false positives).
error: signed integer variable used as array index (-Wsigned-index)
`int i`: should be `size_t i`
error: comparison of signed to unsigned integer (-Wsigned-comparison)
`i < len`: `i` has type `int`; `len` has type `size_t`
Scare quotes on "intent", because I don't write (that particular kind of) obviously broken code in the first place, but the upshot is that i and len both need to be size_t, and anything that lets the compiler obscure that is de facto bad.Edit: previously assumed len was wrong too, but on rereading, you were implying len was already the correct type, and only i was wrong.
(Also, the word "variable" is significant; IIRC something like 5+bytecode_next_char promotes to int, but can't actually be negative[1] so shouldn't generate a warning.)
1: Assuming the compiler is configured correctly, ie char is unsigned.
That seems suspect to me but I understand why they'd do it.
No, they should not. Integer overflow is in most cases logic error, like division by zero or NULL pointer dereference, so it should stay undefined.
"NOTE: Possible undefined behavior ranges from ignoring the situation completely with unpredictable results, to behaving during translation or program execution in a documented manner characteristic of the environment (...), to terminating a translation or execution (...)."
> which is more like "do whatever the sensible thing would be on the target platform"
That is a meaning of 'implementation-defined' in C parlance.
Edit: Ugh, it needs javascript to render simple HTML.
https://icube-forge.unistra.fr/icps/c23-library/-/blob/main/...
Raw Markdown, for no-JS readers:
https://icube-forge.unistra.fr/icps/c23-library/-/raw/main/R...
That was already portable between 16 bit, 32 bit, 64 bit etc. Why is it that just because the compiler supports 128 bit or 256 bit integers that compiling in such a mode doesn't correspondingly update "[u]intmax_t"?
The linked page says they 'cannot be "extended integer types" in the sense of C17', but that printf() and scanf() should still support these?
typedef int intpremium_mediocre_t;First, please stop breaking the site guidelines. You've repeatedly posted unsubstantive/flamebait comments—which as you know, that is not what HN is for. In particular, we ban accounts that do abusive things like https://news.ycombinator.com/item?id=33617059.
Second, please stop routinely creating accounts. As the guidelines say, Throwaway accounts are ok for sensitive information, but please don't create accounts routinely. HN is a community—users should have an identity that others can relate to.
Ironically, this is actually one of the only[0] legitimate uses for standard[1] PRI* macros, since that could expand to whichever of "%llu", "%w128u", etc, was appropriate to the caller.
0: And I'm not sure "one of" is actually needed.
1: as opposed to nonstandard ones like PRIu_xlib_atom or the like
2: Depending on how you define "single", it competes with (x*(uintmax_t)y)>>UINTPTR_BIT, but that's not actually reliable since uintmax_t isn't (IIRC) guaranteed to be larger than uintptr_t.
intmax_199409L_t for C89's 1994 amendment, etc.
And now code written to conform with the intent and wording of C99 will need to change?
If you have an interface that has a function that e.g. takes an intmax_t parameter (and even the C standard has those, e.g. imaxabs()), increasing intmax_t size (ABI change) would break existing callers.
So you can only change intmax_t size if you do not care about ABI stability.
Anyway you can (and should!) use -fvisibility=hidden and add __attribute__((__visibility__("default"))) to public symbols when writing a C library. It will make calls between non-visible symbols faster because the compiler doesn't have to generate code to handle ELF symbol interposition.
That said, there are effectively two standard compiler interfaces: MSVC and GCC and everyone else emulates one or both of these.
Similarily, symbol visibility is not something that the C standard cares about because it doesn't even care about libraries in the first place. Again, for all platforms that have symbol visibility (e.g. PE and ELF based ones, although the details differ) there is a defacto standard for the compiler flags and attributes to control the visilibity: the MSVC and GCC extensions.
Finally! My joy knows no bounds.
Much of the time, neither does a C compiler.
They seem pretty harmless, are they difficult to implement?
#include <stdio.h>
int main() {
puts(“What??(Really.)”);
}
Try running that. Only change the quote marks to the ones on a normal keyboard (I’m on my phone.)Edit: Oh, and remember to compile with -std=c11 e.g. gnu11 has trigraphs disabled by default.
printf("Eh???/n");
Which is the same as: printf("Eh?\n");
I.e. the "??/" supplies the "\" to the "n", making it a "\n". I think it's good that they're gone (they were there to support some ancient non-ASCII systems). But it's worth mentioning that the "horrifying" semantics are there so that these trigraphs combine with <s>digraphs</s> (edit: "escape sequences", see downthread) in this way.Otherwise you'd just get:
printf("Eh?\\n");I.e. that the trigraphs need to be parsed like that both because they're replacements for basic syntax like "{" and "}", but also because they need to be considered as if though the parser saw a "\" in a constant, as they next character may be e.g. "n", forming a "\n", not "\\n".
So, I tried it with gcc and clang, and neither did anything strange.
Also tried -std=c99 and -std=c89. All behaved the same.
I thought maybe my distro had some default option configured to turn of trigraphs/digraphs, but --verbose didn't seem to show anything suspicious.
$ cat garbage.c
# include<stdio.h>
int main()
{
puts("What?? (really)");
return 0;
}
$ clang -std=c11 -o garbage garbage.c
$ ./garbage
What?? (really)
$ clang -std=c99 -o garbage garbage.c
$ ./garbage
What?? (really)
$ man gcc
$ clang -std=c89 -o garbage garbage.c
$ ./garbage
What?? (really)
$ gcc -std=c11 -o garbage garbage.c
$ ./garbage
What?? (really)
$ gcc -std=c99 -o garbage garbage.c
$ ./garbage
What?? (really)
$ gcc -std=c89 -o garbage garbage.c
$ ./garbage
What?? (really)
$It was the whole change of management that somehow made them backtrack on that decision.
> Extended integer types may be wider than intmax_t